Every model's token market has its own structure, its own spread, and its own traps, and comparing them from provider rate cards means opening thirty tabs. This page is the one table: official rates, cheapest available rates, cache terms, and market structure for the major Chinese and open-weight flagships, drawn from the Mercatus per-model pricing pages and updated as they update.
The master table
All prices in USD per million tokens, as of September 2, 2026:
| Model | Official input / output | Official cached input | Cheapest input / output | Hosts | Spread | Context |
|---|---|---|---|---|---|---|
| DeepSeek V4 Flash | $0.22 / $0.66 off-peak ($0.44 / $1.32 peak) | $0.007–$0.014 | $0.0679 / $0.168 (DigitalOcean) | 17 | 6.5x | 1M |
| DeepSeek V3.2 | Retired from official API (was $0.28 / $0.42) | n/a | ~$0.20 / $0.32 | ~10 | ~3x | 163K |
| DeepSeek V4 Pro | $0.66 / $1.98 off-peak ($1.32 / $3.96 peak) | $0.022–$0.044 | $0.66 (official off-peak) / $1.74 (DigitalOcean) | 16 | ~3x | 1M |
| GLM-5.2 | $1.40 / $4.40 | $0.26 | $0.402 / $1.263 | 21 | 3.5x | 1M |
| Qwen3.8 Max | $2.00 / $6.00 | $0.25 | official (single seller) | 1 | — | — |
| Kimi K3 | $3.00 / $15.00 | $0.30 | $2.80 / $14.00 | 11 | 1.23x | 1M |
Full breakdowns: DeepSeek V4, DeepSeek V3.2, DeepSeek V4 Pro, GLM-5.2, and Kimi K3.
Notes that matter: DeepSeek bills by time of day since August 16, with off-peak at half of peak (the schedule); its third parties repriced within three weeks, upward on V4 Pro and not at all on V4 Flash, which left the official API as the second most expensive Flash host. Every DeepSeek model and host is on DeepSeek API pricing. Qwen3.8 Max is the one model here served only by its developer, so it has no spread to compare. Cheapest-host figures are list prices; verify precision, context limits, and rate limits before relying on any discount tier.
The same workload on every model
A production assistant doing 10M input and 1M output tokens per day, official rates, no caching:
| Model | Daily | Annual |
|---|---|---|
| DeepSeek V4 Flash (cheapest host) | $0.85 | $309 |
| DeepSeek V4 Flash (off-peak) | $2.86 | $1,044 |
| DeepSeek V3.2 (cheapest host) | $2.32 | $847 |
| GLM-5.2 (cheapest host) | $5.28 | $1,927 |
| DeepSeek V4 Pro (off-peak) | $8.58 | $3,132 |
| GLM-5.2 (official) | $18.40 | $6,716 |
| Qwen3.8 Max | $26.00 | $9,490 |
| Kimi K3 | $45.00 | $16,425 |
The ladder runs 16x from bottom to top at official rates, and over 50x once the cheapest third-party host for V4 Flash is on it. That is the most important number on this page: choosing the model is a 16x decision, choosing a provider within a model is a 1.2x to 3.5x decision, and scheduling around peak hours on DeepSeek is a 2x decision. Optimize in that order.
What the spreads tell you
The spread column is a market-maturity gauge, and this summer provided a live demonstration of every state:
- GLM-5.2 (3.5x, recently 30x): a commodity-priced model whose launch chaos, a $0.07 promotional tier and a $2.10 premium tier, collapsed into two bands within two weeks. Wide spreads on young models are promotional and temporary, in both directions.
- DeepSeek V4 Pro (~3x) and V4 Flash (6.5x): a market that re-sorted in three weeks. Third parties followed the official increase upward on Pro and ignored it on Flash, where six hosts still sell at DeepSeek's old $0.14 / $0.28 and the official API is now the second most expensive host of its own model.
- Kimi K3 (1.23x): a premium model whose buyers choose on quality, so hosts price at list and compete on latency instead. The tightest token market we track.
- Qwen3.8 Max (no spread): single-seller pricing, which is simply a posted price with no market check at all.
The pattern underneath all four, why identical tokens price differently and what forces convergence, is in Why Token Prices Differ.
The cache dimension
Cache terms vary more than list prices and routinely decide the real bill:
| Model | Cached input as % of list | The implication |
|---|---|---|
| DeepSeek V4 (Pro and Flash) | ~3% | Deepest discount; cache-heavy workloads live here |
| DeepSeek V3.2 | 10% | Strong; historically flipped provider rankings |
| Kimi K3 | 10% | The main cost lever on an expensive model |
| GLM-5.2 | 15–19% | Meaningful, varies by host |
| Qwen3.8 Max | 12.5% | Single seller, single term |
A workload with 60 percent prompt reuse can cut its bill 40 to 60 percent on any of these models, which is more than most provider switches save. Compute your blended rate, input, output, and cache in your real proportions, before comparing anything; that is the logic behind the Standard Blended Price on the Mercatus Token Index, which tracks these models daily with usage weighting.
How to choose, in order
Model first. The 16x ladder dwarfs every other decision. Benchmark the cheapest model that might pass your evals before paying for the premium tier.
Then cache strategy. Structure prompts for reuse; the discounts run 81 to 97 percent.
Then schedule, if you're on DeepSeek: off-peak is half price, here's the clock.
Then provider, with verification: precision, context caps, rate limits, and the knowledge that discount tiers can vanish overnight.
Then lock what you can't afford to have move. August saw a 12x increase on ten days' notice and a 10x discount-tier snapback; forward contracts exist because posted prices are not commitments.
Frequently asked questions
What is the cheapest LLM API in 2026?
Among the models tracked here, DeepSeek V4 Flash at $0.22 input and $0.66 output per million tokens off-peak, with a 1M context window. On a standard 10M/1M daily workload it runs about $1,044 a year at official rates. Through third-party hosts the same model runs as low as $0.0679 / $0.168, about $309 a year on that workload.
Which model is most expensive?
Kimi K3 at $3.00 input and $15.00 output, roughly 16x the Flash workload cost. It is also the strongest of these models on public benchmarks; the premium buys quality.
Why do these prices change so often?
They are posted prices, set unilaterally by sellers and changeable by announcement. In August 2026 alone, DeepSeek raised rates up to 12x and GLM-5.2's cheapest tier repriced 10x. This page and the per-model pages carry last-verified dates for exactly that reason.
What's the difference between official and cheapest pricing?
The developer's API is one host among several for open-weight models. Third parties sometimes undercut it (GLM, K3, V4 Flash at every hour, V4 Pro at peak), sometimes charge more, and the relationship flips model by model and month by month.
How should I compare prices across models?
On your own token mix: input, output, and cached shares in your real proportions, at the hours you actually run. A single blended number per model, the way the Token Index computes its Standard Blended Price, is the only apples-to-apples comparison.
Methodology
All figures from the Mercatus per-model pricing pages, each carrying its own verification date, compiled from official API documentation and public pricing trackers. DeepSeek rates reflect the August 16, 2026 peak/off-peak schedule; DeepSeek third-party rates from OpenRouter, September 2, 2026. Qwen3.8 Max from its official launch pricing, August 2026. Workload examples use stated volumes at official rates unless noted. This page updates as the underlying model pages update. Last verified: 2026-09-02.
