LLM API Price Index

One report on what the whole market costs, computed from every rate we track rather than summarised from memory. Free to quote and cite; the underlying dataset is published openly.

Across 60 models from 17 providers, input rates span $0.03 to $30.00 per 1M tokens: a 1000x spread. Output spans $0.08 to $180.00, a 2250x spread. The median model sits at $1.70 per 1M blended tokens.

Verified 2026-08-16.

Models tracked
60across 17 providers
Input spread
1000xcheapest to most expensive
Median blended rate
$1.70per 1M tokens, 3:1 mix
Median output premium
4.0xoutput vs input
Offer prompt caching
27median 90% discount
Offer batch pricing
18almost always 50% off

How wide is the spread, really?

Wide enough that "what does an LLM API cost" has no useful answer without naming a model. Qwen3.7 Flash charges $0.03 per 1M input tokens; GPT-5.5 Pro charges $30.00. On the same million tokens of input, that is the difference between a rounding error and a real line item.

The distribution is not even, either. The bottom decile of models sits at or below $0.150 per 1M blended tokens, the median at $1.70, and the top decile at or above $10.00. Most of the models are clustered in the cheap half; a handful of frontier tiers stretch the top of the range and pull the mean far above the median, which is why the median is the number quoted here.

The cheapest model from each provider

Cheapest model per provider, input rateUSD per 1M input tokens

Every provider's floor, ranked. This is the most useful single view for a high-volume workload: when the task is simple and the volume is large, the entry tier is usually where the work should run, and the question is which vendor's entry tier to use.

Provider Cheapest model Input / 1M Output / 1M Blended
Alibaba Qwen3.7 Flash $0.03 $0.13 $0.055
Groq Llama 3.1 8B Instant (via Groq) $0.05 $0.08 $0.058
Amazon Nova Micro $0.035 $0.14 $0.061
Cohere Command R7B $0.0375 $0.15 $0.066
Meta Llama 4 Scout $0.08 $0.30 $0.135
StepFun Step 3.5 Flash $0.10 $0.30 $0.150
OpenAI GPT-4.1 nano $0.10 $0.40 $0.175
Google Gemini 2.5 Flash-Lite $0.10 $0.40 $0.175
DeepSeek DeepSeek V4 Flash $0.14 $0.28 $0.175
Perplexity Sonar Small Online $0.20 $0.20 $0.200
Mistral Mistral Small 4 $0.15 $0.60 $0.262
xAI Grok 4.1 Fast $0.20 $0.50 $0.275
Zhipu AI GLM-5 $0.60 $1.92 $0.930
Baidu ERNIE 5.1 $0.56 $2.54 $1.06
Moonshot AI Kimi K2.5 $0.60 $3.00 $1.20
ByteDance Doubao Seed 2.1 Pro $0.85 $4.23 $1.70
Anthropic Claude Haiku 4.5 $1.00 $5.00 $2.00

The flagship tier from each provider

Flagship model per provider, output rateUSD per 1M output tokens

The other end of every lineup. Note how much less the flagship tiers vary than the entry tiers do: competition at the top of the market runs on capability, and competition at the bottom runs on price.

Provider Flagship tracked Input / 1M Output / 1M Blended
OpenAI GPT-5.5 Pro $30.00 $180.00 $67.50
Anthropic Claude Fable 5 $10.00 $50.00 $20.00
Moonshot AI Kimi K3 $3.00 $15.00 $6.00
Perplexity Sonar Pro $3.00 $15.00 $6.00
Amazon Nova Premier 1.0 $2.50 $12.50 $5.00
Google Gemini 3.1 Pro $2.00 $12.00 $4.50
Cohere Command A $2.50 $10.00 $4.38
Mistral Mistral Medium 3.5 $1.50 $7.50 $3.00
xAI Grok 4.6 $2.00 $6.00 $3.00
Alibaba Qwen3.8 Max $2.00 $6.00 $3.00
Zhipu AI GLM-5.2 $1.40 $4.40 $2.15
ByteDance Doubao Seed 2.1 Pro $0.85 $4.23 $1.70
Baidu ERNIE 5.1 $0.56 $2.54 $1.06
Groq Llama 3.3 70B Versatile (via Groq) $0.59 $0.79 $0.640
DeepSeek DeepSeek V4 Pro $0.435 $0.87 $0.544
Meta Llama 4 Maverick $0.20 $0.60 $0.300
StepFun Step 3.5 Flash $0.10 $0.30 $0.150

The output premium

Output costs more than input on almost every model tracked, at a median of 4.0x. The reason is mechanical rather than commercial: generating each output token needs its own forward pass through the model, while an entire prompt can be processed in parallel.

The practical consequence is that your input-to-output ratio, not just your total token count, decides which model is cheapest for you. A retrieval-heavy workload that reads 20,000 tokens and writes 300 is priced almost entirely on the input rate. A drafting tool that reads 500 and writes 2,000 is priced almost entirely on the output rate, where the spread between models is 2250x. Ranking models on input price alone, which most comparison tables do, is misleading for the second case.

The exceptions are worth knowing about: Sonar Huge Online, Sonar, Sonar Small Online charge the same rate for input and output. For generation-heavy work that is a structural advantage, though on Perplexity's Sonar tiers it comes alongside a per-request search fee that changes the arithmetic completely.

Discounts are where the real money is

Switching models is the obvious lever and often the wrong one to reach for first. 27 of the 60 models tracked here publish a cached-input rate at a median discount of about 90%, and 18 publish a batch discount, almost always a flat 50% off both input and output.

A workload that is both cache-friendly and latency-tolerant can therefore cut its bill substantially without changing model at all, and without any of the evaluation risk that a model switch carries. Size the two levers with the prompt caching and batch API calculators before considering a migration.

The geography of the price floor

The gap between Chinese and Western providers is one of the clearest patterns in the data. The median blended rate across the 14 models tracked from DeepSeek, Alibaba, Zhipu, Moonshot, ByteDance, Baidu and StepFun is $1.20 per 1M tokens, against $2.75 for the 46 Western models: 2.3x apart at the median.

Two caveats on reading too much into that. Rates from ByteDance and Baidu are published in yuan and converted here, so they move with the exchange rate rather than sitting still. And price is only one axis of the decision: data residency, latency from your region, and API reliability do not appear on a rate card but decide whether a provider is usable at all.

Full ranking, cheapest first

# Model Provider Input / 1M Output / 1M Blended
1 Qwen3.7 Flash Alibaba $0.03 $0.13 $0.055
2 Llama 3.1 8B Instant (via Groq) Groq $0.05 $0.08 $0.058
3 Nova Micro Amazon $0.035 $0.14 $0.061
4 Command R7B Cohere $0.0375 $0.15 $0.066
5 Nova Lite 1.0 Amazon $0.06 $0.24 $0.105
6 Llama 4 Scout Meta $0.08 $0.30 $0.135
7 Step 3.5 Flash StepFun $0.10 $0.30 $0.150
8 GPT-4.1 nano OpenAI $0.10 $0.40 $0.175
9 Gemini 2.5 Flash-Lite Google $0.10 $0.40 $0.175
10 DeepSeek V4 Flash DeepSeek $0.14 $0.28 $0.175
11 Sonar Small Online Perplexity $0.20 $0.20 $0.200
12 Mistral Small 4 Mistral $0.15 $0.60 $0.262
13 Grok 4.1 Fast xAI $0.20 $0.50 $0.275
14 Llama 4 Maverick Meta $0.20 $0.60 $0.300
15 GPT-5.6 Luna OpenAI $0.20 $1.20 $0.450
16 GPT-5.4 nano OpenAI $0.20 $1.25 $0.463
17 DeepSeek V4 Pro DeepSeek $0.435 $0.87 $0.544
18 Llama 3.3 70B Versatile (via Groq) Groq $0.59 $0.79 $0.640
19 Mistral Large 3 Mistral $0.50 $1.50 $0.750
20 Gemini 3.5 Flash-Lite Google $0.30 $2.50 $0.850
21 GLM-5 Zhipu AI $0.60 $1.92 $0.930
22 GLM-4.6 Zhipu AI $0.60 $2.20 $1.00
23 Sonar Perplexity $1.00 $1.00 $1.00
24 ERNIE 5.1 Baidu $0.56 $2.54 $1.06
25 Kimi K2.5 Moonshot AI $0.60 $3.00 $1.20
26 Grok Build 0.1 xAI $1.00 $2.00 $1.25
27 Nova Pro 1.0 Amazon $0.80 $3.20 $1.40
28 Gemini 3.7 Flash Google $0.75 $3.75 $1.50
29 Grok 4.3 xAI $1.25 $2.50 $1.56
30 GPT-5.4 mini OpenAI $0.75 $4.50 $1.69
31 Doubao Seed 2.1 Pro ByteDance $0.85 $4.23 $1.70
32 Kimi K2.6 Moonshot AI $0.95 $4.00 $1.71
33 Qwen3.7 Max Alibaba $1.25 $3.75 $1.88
34 Claude Haiku 4.5 Anthropic $1.00 $5.00 $2.00
35 GLM-5.2 Zhipu AI $1.40 $4.40 $2.15
36 Magistral Medium Mistral $2.00 $5.00 $2.75
37 Gemini 3.6 Flash Google $1.50 $7.50 $3.00
38 Mistral Medium 3.5 Mistral $1.50 $7.50 $3.00
39 Grok 4.6 xAI $2.00 $6.00 $3.00
40 Grok 4.5 xAI $2.00 $6.00 $3.00
41 Qwen3.8 Max Alibaba $2.00 $6.00 $3.00
42 Gemini 3.5 Flash Google $1.50 $9.00 $3.38
43 Claude Sonnet 5 Anthropic $2.00 $10.00 $4.00
44 Command A Cohere $2.50 $10.00 $4.38
45 Command R+ Cohere $2.50 $10.00 $4.38
46 GPT-5.6 Terra OpenAI $2.00 $12.00 $4.50
47 Gemini 3.1 Pro Google $2.00 $12.00 $4.50
48 Gemini 3 Pro Google $2.00 $12.00 $4.50
49 Nova Premier 1.0 Amazon $2.50 $12.50 $5.00
50 Sonar Huge Online Perplexity $5.00 $5.00 $5.00
51 GPT-5.4 OpenAI $2.50 $15.00 $5.63
52 Claude Sonnet 4.6 Anthropic $3.00 $15.00 $6.00
53 Kimi K3 Moonshot AI $3.00 $15.00 $6.00
54 Sonar Pro Perplexity $3.00 $15.00 $6.00
55 Claude Opus 5 Anthropic $5.00 $25.00 $10.00
56 Claude Opus 4.8 Anthropic $5.00 $25.00 $10.00
57 GPT-5.6 Sol OpenAI $5.00 $30.00 $11.25
58 GPT-5.5 OpenAI $5.00 $30.00 $11.25
59 Claude Fable 5 Anthropic $10.00 $50.00 $20.00
60 GPT-5.5 Pro OpenAI $30.00 $180.00 $67.50

Method

Every figure on this page is computed at build time from the same dataset that drives the rest of the site, so nothing here can drift out of step with the model pages. Rates are standard public list prices in USD per 1M tokens, excluding enterprise and committed-use discounts, which are not public and not comparable. The blended rate assumes three input tokens per output token. Full sourcing is on the methodology page, and the raw data is at /data/.

LLM Cost Lab. "LLM API Price Index." Verified 2026-08-16. https://llmcostlab.com/price-index/

Price index questions

What is the average cost of an LLM API in 2026?

There is no meaningful average, because the distribution is not close to normal: the spread between cheapest and most expensive input rate is 1000x. The median is more useful, and it sits at $1.70 per 1M blended tokens across the 60 models tracked here. Half the market is cheaper than that, half more expensive.

How much more do output tokens cost than input tokens?

The median model charges 4.0x more for output than input. The steepest is Gemini 3.5 Flash-Lite at 8.3x. 3 models charge the same for both, all of them Perplexity Sonar tiers. The reason is mechanical: each output token requires its own forward pass, while input tokens are processed in parallel.

Are Chinese LLM providers cheaper than US ones?

Materially, yes. The median blended rate across the 14 models tracked from DeepSeek, Alibaba, Zhipu, Moonshot, ByteDance, Baidu and StepFun is $1.20 per 1M tokens, against $2.75 for the 46 models from Western providers: roughly 2.3x apart at the median.

How many LLM providers offer prompt caching?

27 of 60 tracked models publish a discounted cached-input rate, at a median discount of about 90% off standard input. 18 publish a batch discount, almost always a flat 50%.

Is LLM pricing going up or down?

Down, and quickly, at the low and mid tiers. OpenAI cut GPT-5.6 Luna by roughly 80% and Terra by 20% on a single day in July 2026, and new Flash-class releases keep arriving at introductory rates below the models they replace. Frontier pricing has been far stickier: the top of the market has not moved much, so the spread between cheapest and most expensive keeps widening.

Go deeper