LLM API Price Index
One report on what the whole market costs, computed from every rate we track rather than summarised from memory. Free to quote and cite; the underlying dataset is published openly.
Across 60 models from 17 providers, input rates span $0.03 to $30.00 per 1M tokens: a 1000x spread. Output spans $0.08 to $180.00, a 2250x spread. The median model sits at $1.70 per 1M blended tokens.
Verified 2026-08-16.
- Models tracked
- 60across 17 providers
- Input spread
- 1000xcheapest to most expensive
- Median blended rate
- $1.70per 1M tokens, 3:1 mix
- Median output premium
- 4.0xoutput vs input
- Offer prompt caching
- 27median 90% discount
- Offer batch pricing
- 18almost always 50% off
How wide is the spread, really?
Wide enough that "what does an LLM API cost" has no useful answer without naming a model. Qwen3.7 Flash charges $0.03 per 1M input tokens; GPT-5.5 Pro charges $30.00. On the same million tokens of input, that is the difference between a rounding error and a real line item.
The distribution is not even, either. The bottom decile of models sits at or below $0.150 per 1M blended tokens, the median at $1.70, and the top decile at or above $10.00. Most of the models are clustered in the cheap half; a handful of frontier tiers stretch the top of the range and pull the mean far above the median, which is why the median is the number quoted here.
The cheapest model from each provider
Every provider's floor, ranked. This is the most useful single view for a high-volume workload: when the task is simple and the volume is large, the entry tier is usually where the work should run, and the question is which vendor's entry tier to use.
| Provider | Cheapest model | Input / 1M | Output / 1M | Blended |
|---|---|---|---|---|
| Alibaba | Qwen3.7 Flash | $0.03 | $0.13 | $0.055 |
| Groq | Llama 3.1 8B Instant (via Groq) | $0.05 | $0.08 | $0.058 |
| Amazon | Nova Micro | $0.035 | $0.14 | $0.061 |
| Cohere | Command R7B | $0.0375 | $0.15 | $0.066 |
| Meta | Llama 4 Scout | $0.08 | $0.30 | $0.135 |
| StepFun | Step 3.5 Flash | $0.10 | $0.30 | $0.150 |
| OpenAI | GPT-4.1 nano | $0.10 | $0.40 | $0.175 |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | $0.175 | |
| DeepSeek | DeepSeek V4 Flash | $0.14 | $0.28 | $0.175 |
| Perplexity | Sonar Small Online | $0.20 | $0.20 | $0.200 |
| Mistral | Mistral Small 4 | $0.15 | $0.60 | $0.262 |
| xAI | Grok 4.1 Fast | $0.20 | $0.50 | $0.275 |
| Zhipu AI | GLM-5 | $0.60 | $1.92 | $0.930 |
| Baidu | ERNIE 5.1 | $0.56 | $2.54 | $1.06 |
| Moonshot AI | Kimi K2.5 | $0.60 | $3.00 | $1.20 |
| ByteDance | Doubao Seed 2.1 Pro | $0.85 | $4.23 | $1.70 |
| Anthropic | Claude Haiku 4.5 | $1.00 | $5.00 | $2.00 |
The flagship tier from each provider
The other end of every lineup. Note how much less the flagship tiers vary than the entry tiers do: competition at the top of the market runs on capability, and competition at the bottom runs on price.
| Provider | Flagship tracked | Input / 1M | Output / 1M | Blended |
|---|---|---|---|---|
| OpenAI | GPT-5.5 Pro | $30.00 | $180.00 | $67.50 |
| Anthropic | Claude Fable 5 | $10.00 | $50.00 | $20.00 |
| Moonshot AI | Kimi K3 | $3.00 | $15.00 | $6.00 |
| Perplexity | Sonar Pro | $3.00 | $15.00 | $6.00 |
| Amazon | Nova Premier 1.0 | $2.50 | $12.50 | $5.00 |
| Gemini 3.1 Pro | $2.00 | $12.00 | $4.50 | |
| Cohere | Command A | $2.50 | $10.00 | $4.38 |
| Mistral | Mistral Medium 3.5 | $1.50 | $7.50 | $3.00 |
| xAI | Grok 4.6 | $2.00 | $6.00 | $3.00 |
| Alibaba | Qwen3.8 Max | $2.00 | $6.00 | $3.00 |
| Zhipu AI | GLM-5.2 | $1.40 | $4.40 | $2.15 |
| ByteDance | Doubao Seed 2.1 Pro | $0.85 | $4.23 | $1.70 |
| Baidu | ERNIE 5.1 | $0.56 | $2.54 | $1.06 |
| Groq | Llama 3.3 70B Versatile (via Groq) | $0.59 | $0.79 | $0.640 |
| DeepSeek | DeepSeek V4 Pro | $0.435 | $0.87 | $0.544 |
| Meta | Llama 4 Maverick | $0.20 | $0.60 | $0.300 |
| StepFun | Step 3.5 Flash | $0.10 | $0.30 | $0.150 |
The output premium
Output costs more than input on almost every model tracked, at a median of 4.0x. The reason is mechanical rather than commercial: generating each output token needs its own forward pass through the model, while an entire prompt can be processed in parallel.
The practical consequence is that your input-to-output ratio, not just your total token count, decides which model is cheapest for you. A retrieval-heavy workload that reads 20,000 tokens and writes 300 is priced almost entirely on the input rate. A drafting tool that reads 500 and writes 2,000 is priced almost entirely on the output rate, where the spread between models is 2250x. Ranking models on input price alone, which most comparison tables do, is misleading for the second case.
The exceptions are worth knowing about: Sonar Huge Online, Sonar, Sonar Small Online charge the same rate for input and output. For generation-heavy work that is a structural advantage, though on Perplexity's Sonar tiers it comes alongside a per-request search fee that changes the arithmetic completely.
Discounts are where the real money is
Switching models is the obvious lever and often the wrong one to reach for first. 27 of the 60 models tracked here publish a cached-input rate at a median discount of about 90%, and 18 publish a batch discount, almost always a flat 50% off both input and output.
A workload that is both cache-friendly and latency-tolerant can therefore cut its bill substantially without changing model at all, and without any of the evaluation risk that a model switch carries. Size the two levers with the prompt caching and batch API calculators before considering a migration.
The geography of the price floor
The gap between Chinese and Western providers is one of the clearest patterns in the data. The median blended rate across the 14 models tracked from DeepSeek, Alibaba, Zhipu, Moonshot, ByteDance, Baidu and StepFun is $1.20 per 1M tokens, against $2.75 for the 46 Western models: 2.3x apart at the median.
Two caveats on reading too much into that. Rates from ByteDance and Baidu are published in yuan and converted here, so they move with the exchange rate rather than sitting still. And price is only one axis of the decision: data residency, latency from your region, and API reliability do not appear on a rate card but decide whether a provider is usable at all.
Full ranking, cheapest first
| # | Model | Provider | Input / 1M | Output / 1M | Blended |
|---|---|---|---|---|---|
| 1 | Qwen3.7 Flash | Alibaba | $0.03 | $0.13 | $0.055 |
| 2 | Llama 3.1 8B Instant (via Groq) | Groq | $0.05 | $0.08 | $0.058 |
| 3 | Nova Micro | Amazon | $0.035 | $0.14 | $0.061 |
| 4 | Command R7B | Cohere | $0.0375 | $0.15 | $0.066 |
| 5 | Nova Lite 1.0 | Amazon | $0.06 | $0.24 | $0.105 |
| 6 | Llama 4 Scout | Meta | $0.08 | $0.30 | $0.135 |
| 7 | Step 3.5 Flash | StepFun | $0.10 | $0.30 | $0.150 |
| 8 | GPT-4.1 nano | OpenAI | $0.10 | $0.40 | $0.175 |
| 9 | Gemini 2.5 Flash-Lite | $0.10 | $0.40 | $0.175 | |
| 10 | DeepSeek V4 Flash | DeepSeek | $0.14 | $0.28 | $0.175 |
| 11 | Sonar Small Online | Perplexity | $0.20 | $0.20 | $0.200 |
| 12 | Mistral Small 4 | Mistral | $0.15 | $0.60 | $0.262 |
| 13 | Grok 4.1 Fast | xAI | $0.20 | $0.50 | $0.275 |
| 14 | Llama 4 Maverick | Meta | $0.20 | $0.60 | $0.300 |
| 15 | GPT-5.6 Luna | OpenAI | $0.20 | $1.20 | $0.450 |
| 16 | GPT-5.4 nano | OpenAI | $0.20 | $1.25 | $0.463 |
| 17 | DeepSeek V4 Pro | DeepSeek | $0.435 | $0.87 | $0.544 |
| 18 | Llama 3.3 70B Versatile (via Groq) | Groq | $0.59 | $0.79 | $0.640 |
| 19 | Mistral Large 3 | Mistral | $0.50 | $1.50 | $0.750 |
| 20 | Gemini 3.5 Flash-Lite | $0.30 | $2.50 | $0.850 | |
| 21 | GLM-5 | Zhipu AI | $0.60 | $1.92 | $0.930 |
| 22 | GLM-4.6 | Zhipu AI | $0.60 | $2.20 | $1.00 |
| 23 | Sonar | Perplexity | $1.00 | $1.00 | $1.00 |
| 24 | ERNIE 5.1 | Baidu | $0.56 | $2.54 | $1.06 |
| 25 | Kimi K2.5 | Moonshot AI | $0.60 | $3.00 | $1.20 |
| 26 | Grok Build 0.1 | xAI | $1.00 | $2.00 | $1.25 |
| 27 | Nova Pro 1.0 | Amazon | $0.80 | $3.20 | $1.40 |
| 28 | Gemini 3.7 Flash | $0.75 | $3.75 | $1.50 | |
| 29 | Grok 4.3 | xAI | $1.25 | $2.50 | $1.56 |
| 30 | GPT-5.4 mini | OpenAI | $0.75 | $4.50 | $1.69 |
| 31 | Doubao Seed 2.1 Pro | ByteDance | $0.85 | $4.23 | $1.70 |
| 32 | Kimi K2.6 | Moonshot AI | $0.95 | $4.00 | $1.71 |
| 33 | Qwen3.7 Max | Alibaba | $1.25 | $3.75 | $1.88 |
| 34 | Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | $2.00 |
| 35 | GLM-5.2 | Zhipu AI | $1.40 | $4.40 | $2.15 |
| 36 | Magistral Medium | Mistral | $2.00 | $5.00 | $2.75 |
| 37 | Gemini 3.6 Flash | $1.50 | $7.50 | $3.00 | |
| 38 | Mistral Medium 3.5 | Mistral | $1.50 | $7.50 | $3.00 |
| 39 | Grok 4.6 | xAI | $2.00 | $6.00 | $3.00 |
| 40 | Grok 4.5 | xAI | $2.00 | $6.00 | $3.00 |
| 41 | Qwen3.8 Max | Alibaba | $2.00 | $6.00 | $3.00 |
| 42 | Gemini 3.5 Flash | $1.50 | $9.00 | $3.38 | |
| 43 | Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | $4.00 |
| 44 | Command A | Cohere | $2.50 | $10.00 | $4.38 |
| 45 | Command R+ | Cohere | $2.50 | $10.00 | $4.38 |
| 46 | GPT-5.6 Terra | OpenAI | $2.00 | $12.00 | $4.50 |
| 47 | Gemini 3.1 Pro | $2.00 | $12.00 | $4.50 | |
| 48 | Gemini 3 Pro | $2.00 | $12.00 | $4.50 | |
| 49 | Nova Premier 1.0 | Amazon | $2.50 | $12.50 | $5.00 |
| 50 | Sonar Huge Online | Perplexity | $5.00 | $5.00 | $5.00 |
| 51 | GPT-5.4 | OpenAI | $2.50 | $15.00 | $5.63 |
| 52 | Claude Sonnet 4.6 | Anthropic | $3.00 | $15.00 | $6.00 |
| 53 | Kimi K3 | Moonshot AI | $3.00 | $15.00 | $6.00 |
| 54 | Sonar Pro | Perplexity | $3.00 | $15.00 | $6.00 |
| 55 | Claude Opus 5 | Anthropic | $5.00 | $25.00 | $10.00 |
| 56 | Claude Opus 4.8 | Anthropic | $5.00 | $25.00 | $10.00 |
| 57 | GPT-5.6 Sol | OpenAI | $5.00 | $30.00 | $11.25 |
| 58 | GPT-5.5 | OpenAI | $5.00 | $30.00 | $11.25 |
| 59 | Claude Fable 5 | Anthropic | $10.00 | $50.00 | $20.00 |
| 60 | GPT-5.5 Pro | OpenAI | $30.00 | $180.00 | $67.50 |
Method
Every figure on this page is computed at build time from the same dataset that drives the rest of the site, so nothing here can drift out of step with the model pages. Rates are standard public list prices in USD per 1M tokens, excluding enterprise and committed-use discounts, which are not public and not comparable. The blended rate assumes three input tokens per output token. Full sourcing is on the methodology page, and the raw data is at /data/.
LLM Cost Lab. "LLM API Price Index." Verified 2026-08-16. https://llmcostlab.com/price-index/
Price index questions
What is the average cost of an LLM API in 2026?
There is no meaningful average, because the distribution is not close to normal: the spread between cheapest and most expensive input rate is 1000x. The median is more useful, and it sits at $1.70 per 1M blended tokens across the 60 models tracked here. Half the market is cheaper than that, half more expensive.
How much more do output tokens cost than input tokens?
The median model charges 4.0x more for output than input. The steepest is Gemini 3.5 Flash-Lite at 8.3x. 3 models charge the same for both, all of them Perplexity Sonar tiers. The reason is mechanical: each output token requires its own forward pass, while input tokens are processed in parallel.
Are Chinese LLM providers cheaper than US ones?
Materially, yes. The median blended rate across the 14 models tracked from DeepSeek, Alibaba, Zhipu, Moonshot, ByteDance, Baidu and StepFun is $1.20 per 1M tokens, against $2.75 for the 46 models from Western providers: roughly 2.3x apart at the median.
How many LLM providers offer prompt caching?
27 of 60 tracked models publish a discounted cached-input rate, at a median discount of about 90% off standard input. 18 publish a batch discount, almost always a flat 50%.
Is LLM pricing going up or down?
Down, and quickly, at the low and mid tiers. OpenAI cut GPT-5.6 Luna by roughly 80% and Terra by 20% on a single day in July 2026, and new Flash-class releases keep arriving at introductory rates below the models they replace. Frontier pricing has been far stickier: the top of the market has not moved much, so the spread between cheapest and most expensive keeps widening.