Qwen3.7 Flash API pricing

Qwen3.7 Flash from Alibaba costs $0.03 per 1M input tokens and $0.13 per 1M output tokens. Batch processing takes 50% off both rates.

That puts it first cheapest of the 60 models tracked here on a blended 3:1 input-to-output rate, or $0.055 per 1M blended tokens. Verified 2026-08-16.

Input
$0.03per 1M tokens
Output
$0.13per 1M tokens
Cached input
None not published
Batch discount
50% off input + output
Price rank
#1of 60, cheapest first
Tokens per $1
18,181,818at a 3:1 mix
Worth knowing: The cheapest model tracked anywhere on this site, on input price.

What Qwen3.7 Flash costs on a real workload

Per-million-token rates are hard to reason about, so here is the same model priced against four concrete monthly workloads. Each row uses this model's own input and output rates against a fixed token mix, with no caching or batch discount applied.

Workload Assumption Monthly cost Annual
Support chatbot 50,000 conversations/month at 3K input and 500 output tokens each $7.75 $93.00
RAG document search 200,000 queries/month at 6K retrieved-context input and 400 output tokens $46.40 $556.80
Coding agent 5,000 runs/month at 60K input and 8K output tokens per run $14.20 $170.40
Bulk classification 5,000,000 items/month at 400 input and 20 output tokens each $73.00 $876.00

Your mix will differ. The API cost calculator takes your own token counts, and the cost per request calculator scales a single call up to per-1,000 and per-month totals.

Models priced near Qwen3.7 Flash

If cost is roughly fixed and you are choosing on capability instead, these are the models sitting closest to Qwen3.7 Flash on price.

Model Provider Input / 1M Output / 1M Blended
Llama 3.1 8B Instant (via Groq) Groq $0.05 $0.08 $0.058
Nova Micro Amazon $0.035 $0.14 $0.061
Command R7B Cohere $0.0375 $0.15 $0.066
Nova Lite 1.0 Amazon $0.06 $0.24 $0.105

How to pay less for Qwen3.7 Flash

Batch processing. Alibaba takes 50% off both input and output when the work goes through the async batch API instead of the real-time one. Anything latency-tolerant qualifies: bulk classification, offline summarisation, dataset labelling, nightly enrichment jobs. The batch API savings calculator shows what that is worth against your volume.

About Alibaba

Alibaba's Qwen line spans a flagship Max tier down to a Flash tier that is the cheapest model on this entire site. Rates differ sharply by region.

See every Alibaba model and how the lineup is tiered on the Alibaba pricing page, or put Qwen3.7 Flash against all 60 models from all 17 providers on the comparison table.

Qwen3.7 Flash pricing FAQ

How much does Qwen3.7 Flash cost per 1M tokens?

Qwen3.7 Flash costs $0.03 per 1M input tokens and $0.13 per 1M output tokens. There is no published cached-input rate for this model.

Is Qwen3.7 Flash expensive compared to other models?

It ranks 1 of 60 on a blended rate that weights input and output 3:1, so 0 tracked models are cheaper and 59 are more expensive. That works out to 1.0x the blended rate of Qwen3.7 Flash, the cheapest model tracked here.

What does a real workload cost on Qwen3.7 Flash?

A support chatbot handling 50,000 conversations a month at 3K input and 500 output tokens each comes to about $7.75 a month. A coding agent doing 5,000 runs at 60K input and 8K output per run comes to about $14.20. Run your own numbers in the API cost calculator.

Why is Qwen3.7 Flash output more expensive than input?

Output on this model is 4.3x the input rate. Every output token needs its own forward pass through the model, while input tokens are processed in parallel, so output costs more to serve across essentially every provider. It also means a workload's input:output mix, not just its total token count, drives the bill.

What is a cheaper alternative to Qwen3.7 Flash?

Nothing tracked here is cheaper. Qwen3.7 Flash is already the lowest blended rate of the 60 models on this site.

Keep going

Rates last checked on 2026-08-16 against Alibaba Cloud Model Studio Pricing. These are standard public list prices and exclude enterprise or committed-use discounts. Use them to plan, not to invoice: see the methodology page for how verification works.