Every model we track, one page each

60 models from 17 providers, each with a dedicated page carrying its input, output, cached-input and batch rates, four worked monthly cost scenarios, and a ranked list of cheaper alternatives. Prices span $0.03 to $30.00 per 1M input tokens, a 1000x spread. Verified 2026-08-16.

See them all ranked in one table Read the price index Download the dataset

Cheapest models tracked

The ten lowest blended rates on the site. Cheap does not mean interchangeable: most of these are small models built for classification, routing, and short completions rather than for reasoning.

# Model Provider Input / 1M Output / 1M Blended
1 Qwen3.7 Flash Alibaba $0.03 $0.13 $0.055
2 Llama 3.1 8B Instant (via Groq) Groq $0.05 $0.08 $0.058
3 Nova Micro Amazon $0.035 $0.14 $0.061
4 Command R7B Cohere $0.0375 $0.15 $0.066
5 Nova Lite 1.0 Amazon $0.06 $0.24 $0.105
6 Llama 4 Scout Meta $0.08 $0.30 $0.135
7 Step 3.5 Flash StepFun $0.10 $0.30 $0.150
8 GPT-4.1 nano OpenAI $0.10 $0.40 $0.175
9 Gemini 2.5 Flash-Lite Google $0.10 $0.40 $0.175
10 DeepSeek V4 Flash DeepSeek $0.14 $0.28 $0.175

All models by provider

Within each provider the models are ordered cheapest first on the blended rate, which usually but not always matches the vendor's own tier naming.

OpenAI 9 models

GPT-4.1 nano · $0.10 / $0.40 GPT-5.6 Luna · $0.20 / $1.20 GPT-5.4 nano · $0.20 / $1.25 GPT-5.4 mini · $0.75 / $4.50 GPT-5.6 Terra · $2.00 / $12.00 GPT-5.4 · $2.50 / $15.00 GPT-5.6 Sol · $5.00 / $30.00 GPT-5.5 · $5.00 / $30.00 GPT-5.5 Pro · $30.00 / $180.00

Google 7 models

Gemini 2.5 Flash-Lite · $0.10 / $0.40 Gemini 3.5 Flash-Lite · $0.30 / $2.50 Gemini 3.7 Flash · $0.75 / $3.75 Gemini 3.6 Flash · $1.50 / $7.50 Gemini 3.5 Flash · $1.50 / $9.00 Gemini 3.1 Pro · $2.00 / $12.00 Gemini 3 Pro · $2.00 / $12.00

Anthropic 6 models

Claude Haiku 4.5 · $1.00 / $5.00 Claude Sonnet 5 · $2.00 / $10.00 Claude Sonnet 4.6 · $3.00 / $15.00 Claude Opus 5 · $5.00 / $25.00 Claude Opus 4.8 · $5.00 / $25.00 Claude Fable 5 · $10.00 / $50.00

xAI 5 models

Grok 4.1 Fast · $0.20 / $0.50 Grok Build 0.1 · $1.00 / $2.00 Grok 4.3 · $1.25 / $2.50 Grok 4.6 · $2.00 / $6.00 Grok 4.5 · $2.00 / $6.00

Amazon 4 models

Nova Micro · $0.035 / $0.14 Nova Lite 1.0 · $0.06 / $0.24 Nova Pro 1.0 · $0.80 / $3.20 Nova Premier 1.0 · $2.50 / $12.50

Mistral 4 models

Mistral Small 4 · $0.15 / $0.60 Mistral Large 3 · $0.50 / $1.50 Magistral Medium · $2.00 / $5.00 Mistral Medium 3.5 · $1.50 / $7.50

Perplexity 4 models

Sonar Small Online · $0.20 / $0.20 Sonar · $1.00 / $1.00 Sonar Huge Online · $5.00 / $5.00 Sonar Pro · $3.00 / $15.00

Alibaba 3 models

Qwen3.7 Flash · $0.03 / $0.13 Qwen3.7 Max · $1.25 / $3.75 Qwen3.8 Max · $2.00 / $6.00

Cohere 3 models

Command R7B · $0.0375 / $0.15 Command A · $2.50 / $10.00 Command R+ · $2.50 / $10.00

Moonshot AI 3 models

Kimi K2.5 · $0.60 / $3.00 Kimi K2.6 · $0.95 / $4.00 Kimi K3 · $3.00 / $15.00

Zhipu AI 3 models

GLM-5 · $0.60 / $1.92 GLM-4.6 · $0.60 / $2.20 GLM-5.2 · $1.40 / $4.40

DeepSeek 2 models

DeepSeek V4 Flash · $0.14 / $0.28 DeepSeek V4 Pro · $0.435 / $0.87

Groq 2 models

Llama 3.1 8B Instant (via Groq) · $0.05 / $0.08 Llama 3.3 70B Versatile (via Groq) · $0.59 / $0.79

Meta 2 models

Llama 4 Scout · $0.08 / $0.30 Llama 4 Maverick · $0.20 / $0.60

Baidu 1 model

ERNIE 5.1 · $0.56 / $2.54

ByteDance 1 model

Doubao Seed 2.1 Pro · $0.85 / $4.23

StepFun 1 model

Step 3.5 Flash · $0.10 / $0.30

Questions about the model directory

How many LLM models does this site track?

60 models across 17 providers, each with its own page showing input, output, cached-input and batch rates plus worked cost examples. The full list is on this page and the ranked table is on the comparison page.

What is the cheapest LLM API model?

Qwen3.7 Flash from Alibaba, at $0.03 per 1M input and $0.13 per 1M output. On a blended 3:1 input-to-output rate it is 1227x cheaper than GPT-5.5 Pro, the most expensive model tracked here.

What is the blended rate used to rank these models?

A single figure assuming three input tokens for every output token, which is closer to real chat and retrieval traffic than either rate alone. Ranking on input price only flatters models with cheap input and expensive output, and that gap can be six-fold or more.

How do I pick a model on cost?

Start from the workload, not the price list. Work out your input:output token ratio and monthly volume, then price the shortlist against it in the API cost calculator. A model with a low input rate can still be the expensive choice for a generation-heavy workload.

Related pages

All rates last checked on 2026-08-16. Standard public list prices only, in USD per 1M tokens. See the methodology page for sourcing.