Gemini 3.5 Flash API pricing

Gemini 3.5 Flash from Google costs $1.50 per 1M input tokens and $9.00 per 1M output tokens. Cached input drops to $0.15 per 1M. Batch processing takes 50% off both rates.

That puts it 42th cheapest of the 60 models tracked here on a blended 3:1 input-to-output rate, or $3.38 per 1M blended tokens. Verified 2026-08-16.

Input
$1.50per 1M tokens
Output
$9.00per 1M tokens
Cached input
$0.15 90% off input
Batch discount
50% off input + output
Price rank
#42of 60, cheapest first
Tokens per $1
296,296at a 3:1 mix

What Gemini 3.5 Flash costs on a real workload

Per-million-token rates are hard to reason about, so here is the same model priced against four concrete monthly workloads. Each row uses this model's own input and output rates against a fixed token mix, with no caching or batch discount applied.

Workload Assumption Monthly cost Annual
Support chatbot 50,000 conversations/month at 3K input and 500 output tokens each $450.00 $5,400
RAG document search 200,000 queries/month at 6K retrieved-context input and 400 output tokens $2,520 $30,240
Coding agent 5,000 runs/month at 60K input and 8K output tokens per run $810.00 $9,720
Bulk classification 5,000,000 items/month at 400 input and 20 output tokens each $3,900 $46,800

Your mix will differ. The API cost calculator takes your own token counts, and the cost per request calculator scales a single call up to per-1,000 and per-month totals.

Cheaper alternatives to Gemini 3.5 Flash

These are the closest cheaper options from other providers, ordered by how near they sit to Gemini 3.5 Flash on the blended rate. Closer is usually a more realistic swap: the further down this list you go, the more capability you are likely trading away.

Model Provider Input / 1M Output / 1M Cheaper by
Qwen3.8 Max Alibaba $2.00 $6.00 1.1x
Grok 4.5 xAI $2.00 $6.00 1.1x
Mistral Medium 3.5 Mistral $1.50 $7.50 1.1x
GLM-5.2 Zhipu AI $1.40 $4.40 1.6x
Claude Haiku 4.5 Anthropic $1.00 $5.00 1.7x

Before switching, price the move properly: the model switching savings calculator puts two models against the same workload and shows the annual difference.

Models priced near Gemini 3.5 Flash

If cost is roughly fixed and you are choosing on capability instead, these are the models sitting closest to Gemini 3.5 Flash on price.

Model Provider Input / 1M Output / 1M Blended
Gemini 3.6 Flash Google $1.50 $7.50 $3.00
Mistral Medium 3.5 Mistral $1.50 $7.50 $3.00
Grok 4.6 xAI $2.00 $6.00 $3.00
Grok 4.5 xAI $2.00 $6.00 $3.00

How to pay less for Gemini 3.5 Flash

Prompt caching. Cache-hit input tokens bill at $0.15 instead of $1.50, which is 90% off. This matters most for workloads that resend a large, mostly-static prefix on every call: system prompts, tool definitions, retrieved documents, long conversation histories. The saving scales with your cache hit rate, so the prompt caching savings calculator is the honest way to size it rather than assuming a perfect hit rate.

Batch processing. Google takes 50% off both input and output when the work goes through the async batch API instead of the real-time one. Anything latency-tolerant qualifies: bulk classification, offline summarisation, dataset labelling, nightly enrichment jobs. The batch API savings calculator shows what that is worth against your volume.

About Google

Google's Gemini lineup runs from Pro down to Flash and Flash-Lite, with batch discounts across the board and frequent introductory pricing on new Flash releases.

See every Google model and how the lineup is tiered on the Google pricing page, or put Gemini 3.5 Flash against all 60 models from all 17 providers on the comparison table.

Gemini 3.5 Flash pricing FAQ

How much does Gemini 3.5 Flash cost per 1M tokens?

Gemini 3.5 Flash costs $1.50 per 1M input tokens and $9.00 per 1M output tokens. Cached input tokens are cheaper again at $0.15 per 1M, a 90% discount on repeated context.

Is Gemini 3.5 Flash expensive compared to other models?

It ranks 42 of 60 on a blended rate that weights input and output 3:1, so 41 tracked models are cheaper and 18 are more expensive. That works out to 61x the blended rate of Qwen3.7 Flash, the cheapest model tracked here.

What does a real workload cost on Gemini 3.5 Flash?

A support chatbot handling 50,000 conversations a month at 3K input and 500 output tokens each comes to about $450.00 a month. A coding agent doing 5,000 runs at 60K input and 8K output per run comes to about $810.00. Run your own numbers in the API cost calculator.

Why is Gemini 3.5 Flash output more expensive than input?

Output on this model is 6.0x the input rate. Every output token needs its own forward pass through the model, while input tokens are processed in parallel, so output costs more to serve across essentially every provider. It also means a workload's input:output mix, not just its total token count, drives the bill.

What is a cheaper alternative to Gemini 3.5 Flash?

Qwen3.8 Max from Alibaba is the closest cheaper option at $2.00 / $6.00, roughly 1.1x cheaper on a blended basis. Whether it is a real substitute depends on whether your task actually needs the extra capability.

Keep going

Rates last checked on 2026-08-16 against Google AI Pricing. These are standard public list prices and exclude enterprise or committed-use discounts. Use them to plan, not to invoice: see the methodology page for how verification works.