Every model we track, one page each
60 models from 17 providers, each with a dedicated page carrying its input, output, cached-input and batch rates, four worked monthly cost scenarios, and a ranked list of cheaper alternatives. Prices span $0.03 to $30.00 per 1M input tokens, a 1000x spread. Verified 2026-08-16.
Cheapest models tracked
The ten lowest blended rates on the site. Cheap does not mean interchangeable: most of these are small models built for classification, routing, and short completions rather than for reasoning.
| # | Model | Provider | Input / 1M | Output / 1M | Blended |
|---|---|---|---|---|---|
| 1 | Qwen3.7 Flash | Alibaba | $0.03 | $0.13 | $0.055 |
| 2 | Llama 3.1 8B Instant (via Groq) | Groq | $0.05 | $0.08 | $0.058 |
| 3 | Nova Micro | Amazon | $0.035 | $0.14 | $0.061 |
| 4 | Command R7B | Cohere | $0.0375 | $0.15 | $0.066 |
| 5 | Nova Lite 1.0 | Amazon | $0.06 | $0.24 | $0.105 |
| 6 | Llama 4 Scout | Meta | $0.08 | $0.30 | $0.135 |
| 7 | Step 3.5 Flash | StepFun | $0.10 | $0.30 | $0.150 |
| 8 | GPT-4.1 nano | OpenAI | $0.10 | $0.40 | $0.175 |
| 9 | Gemini 2.5 Flash-Lite | $0.10 | $0.40 | $0.175 | |
| 10 | DeepSeek V4 Flash | DeepSeek | $0.14 | $0.28 | $0.175 |
All models by provider
Within each provider the models are ordered cheapest first on the blended rate, which usually but not always matches the vendor's own tier naming.
OpenAI 9 models
Google 7 models
Anthropic 6 models
xAI 5 models
Amazon 4 models
Mistral 4 models
Perplexity 4 models
Alibaba 3 models
Cohere 3 models
Moonshot AI 3 models
Zhipu AI 3 models
DeepSeek 2 models
Groq 2 models
Meta 2 models
Baidu 1 model
ByteDance 1 model
StepFun 1 model
Questions about the model directory
How many LLM models does this site track?
60 models across 17 providers, each with its own page showing input, output, cached-input and batch rates plus worked cost examples. The full list is on this page and the ranked table is on the comparison page.
What is the cheapest LLM API model?
Qwen3.7 Flash from Alibaba, at $0.03 per 1M input and $0.13 per 1M output. On a blended 3:1 input-to-output rate it is 1227x cheaper than GPT-5.5 Pro, the most expensive model tracked here.
What is the blended rate used to rank these models?
A single figure assuming three input tokens for every output token, which is closer to real chat and retrieval traffic than either rate alone. Ranking on input price only flatters models with cheap input and expensive output, and that gap can be six-fold or more.
How do I pick a model on cost?
Start from the workload, not the price list. Work out your input:output token ratio and monthly volume, then price the shortlist against it in the API cost calculator. A model with a low input rate can still be the expensive choice for a generation-heavy workload.