A blended million tokens of US-model inference costs $4.56 today. The same blended million on Chinese models costs $1.12. That 4x gap sounds wide until you see where it was a month ago: the US index has fallen 45 percent in 30 days while the Chinese index barely moved, which means the gap between the two markets compressed from roughly 7x to 4x in a single month.
Those numbers come from the Mercatus Token Price Index, which tracks the blended cost of AI inference across regions daily, weighted by real platform usage. This piece covers what the two curves are doing, why, and what a move this fast means for anyone whose inference bill lives on either side of it.
What the TPI measures
The Token Price Index computes a Standard Blended Price per model, a single cost per million tokens weighting input, cache, and output the way real workloads consume them, then aggregates models into regional benchmarks weighted by actual usage on the Mercatus platform. US-TPI currently blends nine models across Google, Anthropic, and OpenAI; CN-TPI blends seven Chinese models. Weights rebalance daily from a rolling 7-day usage window.
Two properties matter for reading what follows. First, this is a usage-weighted index, not a list-price average: it moves when providers cut prices and when users shift traffic toward cheaper models, and both are real components of what inference actually costs. Second, all prices are quoted in USD per million tokens from published provider rates.
The two curves
| Index | Level (Aug 2026) | 30-day change |
|---|---|---|
| US-TPI (9 models) | $4.56 | −45.0% |
| CN-TPI (7 models) | $1.12 | −3.2% |
| Coding-TPI (8 models) | $8.30 | +4.2% |
Thirty days ago the implied levels were roughly $8.30 for US-TPI and $1.16 for CN-TPI, a ratio of about 7.2x. Today it is 4.1x. Nearly half the premium for US-model inference evaporated in one month.
What's inside the US index
The US-TPI's nine components, with blended prices and current usage weights:
| Model | Provider | SBP $/1M | Weight |
|---|---|---|---|
| Gemini 3 Flash Preview | $1.6250 | 21.4% | |
| Claude Sonnet 4.6 | Anthropic | $9.2250 | 18.5% |
| Gemini 2.5 Flash Lite | $0.2450 | 16.8% | |
| Gemini 2.5 Flash | $1.3350 | 12.8% | |
| gpt-oss-120b | OpenAI | $0.1368 | 12.3% |
| Claude Opus 4.7 | Anthropic | $15.3750 | 6.7% |
| Gemini 3.1 Pro | $6.5125 | 4.5% | |
| GPT-5.4 | OpenAI | $11.8750 | 3.7% |
| Claude Opus 4.6 | Anthropic | $15.3750 | 3.3% |
The composition explains the collapse at a glance. More than 60 percent of US inference usage now runs on models blending under $1.65 per million tokens, while the premium tier ($9 to $15) carries barely a third of the weight it prices for. The index isn't falling because premium models got cheap; it's falling because most tokens stopped being premium tokens.
Why the US index collapsed
The 45 percent decline decomposes into the two forces the index is built to capture.
New frontier models shipped cheaper. The US index's largest weight is now a flash-tier frontier preview model with a blended price around $1.63 per million tokens, carrying over 21 percent of the index. When a frontier lab ships a model at a fraction of the incumbent price and it works, the price cut propagates through every workload that adopts it.
Usage rotated down the price curve. Because the TPI weights by real usage, the index also fell as traffic moved from premium models (blended prices of $9 to $15 per million) toward the new cheap tier. That rotation is not an artifact; it is the market repricing itself. The average dollar spent on US inference buys dramatically more than it did in July.
Both forces point the same direction, and neither is exhausted: the premium models still carry meaningful index weight, which is room for further decline if usage keeps rotating.
Why the Chinese index is holding
CN-TPI at $1.12 moved just 3.2 percent in the same month. Chinese near-frontier models were already priced at open-weight competitive levels; there was far less premium to compress. The GLM-5.2 hosting market, with its discount tiers near $0.07 per million input tokens, shows how much hosting competition has already squeezed that side of the market.
The result is a floor effect. When one side of a gap is already near serving cost, convergence can only come from the expensive side falling, and that is exactly what the data shows: this convergence is a US price collapse, not a Chinese price climb.
The counterpoint: coding inference got more expensive
While US-TPI fell 45 percent, Coding-TPI rose 4.2 percent to $8.30, nearly double the general US index. Coding workloads concentrate on premium models where output quality justifies the rate, and demand for those tokens is strong enough that prices firmed while everything else fell. One market, two directions: commodity-tier inference is deflating fast, and the premium coding tier is holding pricing power. Anyone modeling "token prices only go down" against their coding spend is using the wrong curve.
What a 45 percent monthly move means for buyers
In budget terms: a workload running 1B blended tokens per month on the US mix cost about $8,300 a month in July and costs about $4,560 today, roughly $45,000 a year in savings that arrived without anyone negotiating anything. The same workload on the Chinese mix costs $1,120 a month, but that advantage is now 4x instead of 7x.
If you buy US-model inference: do not sign anything long at last month's prices. A 45 percent monthly decline means any annual commitment priced off a stale quote overpays immediately. Short commitments, frequent re-benchmarking against the index, and pressure on providers to pass through the new frontier pricing.
If you buy Chinese-model inference: your price is stable but your relative advantage is shrinking. The 7x discount that justified cross-region routing complexity last month is 4x today. If the gap keeps closing, the operational case for splitting workloads across regions weakens with it.
If you sell inference: the deflation is not uniform, and the Coding-TPI shows where pricing power still lives. Providers positioned on premium coding workloads are the only segment that raised effective prices this month.
For everyone: a market that moves 45 percent in a month is not a market you can budget against with a posted price. This is the volatility that forward contracts exist to absorb: a price fixed at trade time, settled in actual tokens, regardless of what either index does in between. The mechanics are in How a Token Exchange Works.
The bigger pattern
Every fragmented market compresses eventually: we have documented the same sequence in provider-level token spreads and GPU rentals. What is unusual here is speed. Oil and freight took years to converge after their markets opened; blended AI inference just closed half a regional gap in thirty days. Prices that move this fast make posted-price procurement unmanageable, and they make visible, continuously updated benchmarks the minimum requirement for buying intelligently. That is what the Token Index is for.
Frequently asked questions
How much cheaper is Chinese AI inference than US inference?
About 4x on a blended basis as of August 2026: $1.12 versus $4.56 per blended million tokens. A month earlier the gap was roughly 7x.
Why are US token prices falling?
New frontier models shipped at flash-tier prices, and usage rotated toward them. Both effects compound in a usage-weighted index, producing the 45 percent 30-day decline.
Are Chinese token prices rising?
Over the last 30 days, no: CN-TPI moved just 3.2 percent. Chinese model pricing was already near open-weight competitive levels, which is why the convergence is being driven from the US side.
What is the Token Price Index?
A set of daily benchmark indices (US-TPI, CN-TPI, Coding-TPI) tracking the blended cost of AI inference, computed from published provider prices and weighted by real platform usage. Full methodology on the Token Index page.
Why is coding inference more expensive?
Coding workloads concentrate on premium models whose output quality commands higher rates, and demand is strong enough that Coding-TPI rose 4.2 percent in the same month the general US index fell 45 percent.
Can you hedge against token price moves?
Yes, as of August 2026: forward contracts on the Mercatus Forward Market fix a token price for future delivery, physically settled in tokens. For a market moving 45 percent monthly, that is the difference between a budget and a guess.
Methodology
All index levels and changes from the Mercatus Token Price Index, August 2026. The TPI computes a Standard Blended Price per model (weighted input, cache, and output) and aggregates by real platform usage over a rolling 7-day window, rebalanced daily. 30-day changes reflect both provider price changes and usage-weight rebalancing. Prior-month levels are derived arithmetically from current levels and 30-day changes. Last verified: 2026-08-11.
