Alibaba · Efficient
Qwen 3.7 Flash
Currently the cheapest per-token API of any model we track, at three cents per million input tokens. Worth a look for anything high-volume and simple.
What that means
A worked example
Per-token pricing is hard to feel. Take a moderately chatty production workload: 100 calls a day, roughly 10,000 tokens in and 2,000 out each time. On Qwen 3.7 Flash that comes to about:
$2 / month
Before caching, batching or volume discounts, all of which move this number a long way. Illustrative only.
Strong at
- -Cheapest API pricing tracked
- -Throughput
Typical use
- -Bulk classification
- -Tagging
- -Simple extraction
Same tier
What else to look at
| Model | Developer | Context | In / 1M | Out / 1M |
|---|---|---|---|---|
| Claude Haiku 4.5 | Anthropic | 200K | $1 | $5 |
| GPT-5.6 Luna | OpenAI | 1M | $0.20 | $1.20 |
| DeepSeek V4 Pro | DeepSeek | 1M | $0.43 | $0.87 |
| DeepSeek V4 Flash | DeepSeek | 1M | $0.14 | $0.28 |
| Qwen 3.6 27B | Alibaba | - | - | - |
Verified 2026-08-07.