Gemini 3.1 Pro vs Gemini 3.7 Flash
Gemini 3.7 Flash is the cheaper of the two, at 3.0x less on a blended 3:1 input-to-output rate. Gemini 3.1 Pro runs $2.00 in / $12.00 out per 1M tokens; Gemini 3.7 Flash runs $0.75 in / $3.75 out.
On the heaviest scenario below, that gap is $3,325 a month ($39,900 a year) for identical volume. Verified 2026-08-16.
Side by side
| Gemini 3.1 Pro | Gemini 3.7 Flash | |
|---|---|---|
| Provider | ||
| Input per 1M tokens | $2.00 | $0.75 |
| Output per 1M tokens | $12.00 | $3.75 |
| Cached input per 1M | Not offered | Not offered |
| Batch discount | 50% | 50% |
| Output-to-input ratio | 6.0x | 5.0x |
| Blended per 1M (3:1) | $4.50 | $1.50 |
What each costs on the same workload
Rate cards are hard to compare directly because the input:output mix changes the answer. These four scenarios hold the workload fixed and vary only the model, with no caching or batch discount applied to either side.
| Workload | Gemini 3.1 Pro | Gemini 3.7 Flash | Monthly difference | Cheaper |
|---|---|---|---|---|
| Support chatbot 50,000 conversations/month at 3K input and 500 output tokens each | $600.00 | $206.25 | $393.75 | Gemini 3.7 Flash |
| RAG document search 200,000 queries/month at 6K retrieved-context input and 400 output tokens | $3,360 | $1,200 | $2,160 | Gemini 3.7 Flash |
| Coding agent 5,000 runs/month at 60K input and 8K output tokens per run | $1,080 | $375.00 | $705.00 | Gemini 3.7 Flash |
| Bulk classification 5,000,000 items/month at 400 input and 20 output tokens each | $5,200 | $1,875 | $3,325 | Gemini 3.7 Flash |
Which should you actually pick?
Gemini 3.7 Flash is cheaper on both input and output, so on price alone it wins regardless of your token mix. That does not automatically make it the right call: these are different models with different capability profiles, and paying 3.0x more is justified whenever the more expensive model gets the task right on the first attempt and the cheaper one needs two or three tries, or needs human correction downstream.
Discounts can also flip the arithmetic. Neither model publishes a cached-input rate, so that lever is unavailable on both sides. If the work is latency-tolerant, the batch API savings calculator is worth a look before you decide.
Gemini 3.1 Pro vs Gemini 3.7 Flash: common questions
Is Gemini 3.1 Pro or Gemini 3.7 Flash cheaper?
Gemini 3.7 Flash is cheaper overall, at 3.0x less on a blended 3:1 input-to-output rate ($1.50 vs $4.50 per 1M blended tokens). It is cheaper on both input and output.
What is the price difference between Gemini 3.1 Pro and Gemini 3.7 Flash on a real workload?
On the bulk classification scenario (5,000,000 items/month at 400 input and 20 output tokens each), Gemini 3.1 Pro costs $5,200 a month and Gemini 3.7 Flash costs $1,875. That is a difference of $3,325 a month, or $39,900 a year, for the same volume of work.
Does Gemini 3.1 Pro or Gemini 3.7 Flash have better discounts?
Gemini 3.1 Pro has no published cached-input rate and a 50% batch discount. Gemini 3.7 Flash has no published cached-input rate and a 50% batch discount. On cache-heavy or latency-tolerant workloads these can matter more than the headline rate.
Should I switch from Gemini 3.1 Pro to Gemini 3.7 Flash?
Only if the cheaper model actually does the job. Price is the easy half of the decision; the hard half is whether output quality holds on your task. Run both against a sample of real traffic, then use the model switching savings calculator to put a number on the annual difference before committing.