Two weeks ago, GLM-5.2 had the widest price spread we had ever documented: $0.07 to $2.10 per million input tokens across twenty hosts, a 30x gap. That market just restructured. The discount tier vanished, the premium tier collapsed, and GLM-5.2 now prices between $0.402 and $1.40 across 21 providers, a 3.5x spread. The official Z.AI API never moved: $1.40 input, $4.40 output, through all of it.
This page has the current rates, and the story of the fastest spread compression we have tracked, including what happened to buyers who built on the $0.07 tier. Spoiler: when we published the original version of this page, we wrote that discount-tier pricing "may reflect promotional pricing, loss-leader routing, or quantized serving" and that a workload moved onto it might sit on "a different product with a countdown timer." The countdown ran out in under three weeks.
What GLM-5.2 is, briefly
Released in June 2026, GLM-5.2 is a widely hosted open-weight model with a 1M-token context window, tool calling, and prompt caching, now serving at roughly 110 tokens per second. The full 1M window bills at standard rates with no long-context surcharge, which remains its standout economic feature for long-document and agent workloads. It is also the first model whose forward price was set by a market: the first contract traded on the Mercatus Forward Market, physically settled in tokens.
GLM-5.2 official API pricing
| Token type | Price per 1M tokens (Z.AI official) |
|---|---|
| Input | $1.40 |
| Input, cache hit | $0.26 |
| Output | $4.40 |
Unchanged since launch, through a 95 percent collapse in third-party pricing and the snapback that followed.
GLM-5.2 pricing by provider: all 21 hosts
| Provider | Input $/1M | Output $/1M | Cached input $/1M |
|---|---|---|---|
| FlexAI | $0.402 | $1.263 | $0.06 |
| StreamLake | $0.693 | $2.178 | $0.129 |
| Novita | $0.70 | $2.20 | $0.13 |
| DeepInfra | $0.75 | $2.40 | — |
| Morph | $0.842 | $3.136 | $0.168 |
| Alibaba | $0.966 | $3.036 | $0.193 |
| OpenRouter (routed) | $0.966 | $3.036 | $0.193 |
| SiliconFlow | $1.19 | $3.74 | $0.221 |
| Phala | $1.26 | $3.00 | $0.22 |
| AtlasCloud | $1.26 | $3.96 | $0.234 |
| Parasail | $1.40 | $4.40 | $0.26 |
| Z.AI (official) | $1.40 | $4.40 | $0.26 |
| Friendli | $1.40 | $4.40 | $0.26 |
| Cloudflare | $1.40 | $4.40 | $0.26 |
| Together | $1.40 | $4.40 | — |
| Venice | $1.40 | $4.40 | $0.26 |
| Crusoe | $1.40 | $4.40 | $0.26 |
| BaseTen | $1.40 | $4.40 | $0.14 |
| GMI Cloud | $1.40 | $4.40 | $0.26 |
| Fireworks | $1.40 | $4.40 | $0.14 |
| Mistral | $1.40 | $4.40 | $0.14 |
The market has consolidated into two bands: a competitive tier from $0.40 to $1.26, and an official-parity cluster of ten hosts at exactly $1.40. The old extremes are gone at both ends: nobody serves at $0.07 anymore, and nobody charges $2.10.
What just happened: anatomy of a spread collapse
The version of this page published August 10 documented a 30x spread and a cheapest price of $0.07, itself the result of a 95 percent price collapse in the 90 days after launch. Within two weeks:
- The discount tier disappeared. StreamLake repriced from $0.07 to $0.693, roughly 10x overnight. Novita followed the same path. OpenRouter's routed rate, which passed through the cheapest host, went from $0.07 to $0.966, a 14x move for anyone riding the router.
- The premium tier folded into parity. Cloudflare and Fireworks dropped from $2.10 to exactly $1.40, matching official list.
- A new floor emerged. FlexAI at $0.402 is the cheapest host, at 5.7x the old floor.
Net effect: the cheapest available price rose 5.7x while the spread compressed from 30x to 3.5x. Anyone who moved production onto the $0.07 tier without a fallback saw their inference bill multiply in a fortnight, which is the concrete version of the warning in the original article: a price that looks like an outlier usually is one.
This is also what market maturation looks like, compressed into weeks instead of the years it took DeepSeek V3.2's spread to reach ~3x. Promotional pricing burns off, unsustainable margins correct in both directions, and the surviving spread reflects real cost differences rather than customer-acquisition spend. The pattern, and why posted-price markets keep doing this, is documented in Why Token Prices Differ.
What a real workload pays now
A production assistant doing 10M input and 1M output tokens per day, no caching:
| Provider | Daily | Monthly | Annual |
|---|---|---|---|
| FlexAI (cheapest, $0.402/$1.263) | $5.28 | $158.40 | $1,927 |
| Official Z.AI ($1.40/$4.40) | $18.40 | $552 | $6,716 |
With a 60 percent cache hit rate on input, FlexAI ($0.06 cached) drops to $3.23 a day against the official API's $11.56. The cheapest tier still wins on every mix, but the stakes changed: the gap is now 3.5x instead of 20x, and the events of the last two weeks are the argument for not building a budget on any single host's list price. The same workload cost $0.92 a day on August 10 at a price that no longer exists.
How to choose a GLM-5.2 provider now
- Cost-driven workloads: FlexAI's $0.402 tier, with the standard verification (precision, full 1M context, rate limits), and, after this month, a tested fallback config for when the floor moves again.
- Production workloads valuing stability: the $1.40 official-parity cluster just demonstrated it doesn't move; that stability is what the premium buys.
- Budget certainty: a forward contract remains the only way to actually fix a price in a market that repriced 10x in two weeks; GLM-5.2's forwards trade on the Forward Market with physical settlement.
- Whatever you choose, recheck monthly. This page's own history is the case for it.
Frequently asked questions
How much does GLM-5.2 cost per million tokens?
$1.40 input and $4.40 output on the official Z.AI API, with cached input at $0.26. Third-party hosts range from $0.402 to $1.40 input across 21 providers as of late August 2026.
Who is the cheapest GLM-5.2 provider?
FlexAI, at $0.402 input and $1.263 output, with cached input at $0.06. The previous cheapest tier at $0.07 was repriced roughly 10x in mid-August.
What happened to the $0.07 GLM-5.2 pricing?
It ended. StreamLake and the hosts routing through it repriced to $0.69 and above in mid-August 2026, collapsing the model's 30x price spread to about 3.5x. Discount-tier pricing at that depth was promotional, and it behaved like it.
Has GLM-5.2 gotten cheaper since launch?
Net, yes: the cheapest available price is down about 71 percent from the $1.40 launch rate. But the path was violent: down 95 percent to $0.07, then back up 5.7x to $0.402, all within a quarter.
Does GLM-5.2 charge more for long context?
No. The full 1M-token window bills at standard rates on the official API, with no long-context surcharge.
What is a market-set forward price?
A price for future token delivery set by a trade on an exchange rather than a provider's list. GLM-5.2 was the first model with one, and a spot market that moves 10x in two weeks is the clearest argument for why forwards exist.
Methodology
Provider pricing compiled from published rate cards and public pricing trackers (sourced from OpenRouter and Helicone data), August 24-25, 2026. Historical pricing (the $0.07 floor and 30x spread) as documented in the August 10 version of this page and public tracker history. Forward market facts reflect publicly announced Mercatus contracts. Last verified: 2026-08-25.
