HomeBlogGLM-5.2 API Pricing: $0.07 to $2.10, a 30x Spread
GeneralAug 10, 20267 min read

GLM-5.2 API Pricing: $0.07 to $2.10, a 30x Spread

GLM-5.2 costs $1.40 per million input tokens on the official API, but 20 hosts price it from $0.07 to $2.10. The 30x spread, the 95% collapse, and the first market-set forward price.

M

Mercatus Compute

Author

GLM-5.2 API Pricing: $0.07 to $2.10, a 30x Spread

GLM-5.2 costs $1.40 per million input tokens and $4.40 per million output tokens on the official Z.AI API. Across the twenty providers hosting it, the same model prices anywhere from $0.07 to $2.10 per million input tokens. That is a 30x spread, the widest we have documented for any model, and the cheapest price has fallen 95 percent in roughly 90 days while the official list price has not moved.

GLM-5.2 is also, as of August 2026, the first model whose forward price is set by a market: it is the first contract to trade on the Mercatus Forward Market, with physical settlement in tokens at the end of August. More on what that means below.

This page is the pricing reference: official rates, the full provider table, what a real workload pays with and without caching, how GLM-5.2 compares against DeepSeek V3.2 on cost, and how to read a market this young. For the DeepSeek equivalent, see DeepSeek V3.2 API Pricing.

What GLM-5.2 is, briefly

Released in June 2026, GLM-5.2 is a widely hosted open-weight model with a 1M-token context window, tool calling, and prompt caching, serving at roughly 97 tokens per second with a median time-to-first-token just under 2 seconds. On public benchmark aggregates it sits in the 86th percentile for general intelligence and 68th for coding, which places it in the near-frontier tier where most production open-weight workloads live.

One pricing property stands out before any provider comparison: the full 1M window bills at standard rates with no long-context surcharge. Most models with windows this large charge tiered rates above 128K or 200K. For long-document and agent workloads, that single policy is worth more than most provider discounts.

GLM-5.2 official API pricing

Token typePrice per 1M tokens (Z.AI official)
Input$1.40
Input, cache hit$0.26
Output$4.40

Cache hits bill at roughly a fifth of the standard input rate, and Z.AI is currently waiving cache-storage fees.

GLM-5.2 pricing by provider: all 20 hosts

ProviderInput $/1MOutput $/1MCached input $/1M
StreamLake$0.07$0.22$0.013
OpenRouter (routed)$0.07$0.22$0.013
Novita$0.081$0.255$0.015
DeepInfra$0.75$2.40
GMI Cloud$0.924$2.904$0.172
Alibaba$0.966$3.036$0.193
Morph$1.10$4.10$0.22
Phala$1.134$3.564$0.211
SiliconFlow$1.19$3.74$0.221
AtlasCloud$1.26$3.96$0.234
WandB$1.39$4.40$0.26
Z.AI (official)$1.40$4.40$0.26
Parasail$1.40$4.40$0.26
Friendli$1.40$4.40$0.26
Together$1.40$4.40
Venice$1.40$4.40$0.26
Crusoe$1.40$4.40$0.26
BaseTen$1.40$4.40$0.14
Cloudflare$2.10$6.60$0.21
Fireworks$2.10$6.60$0.21

The market splits into four visible tiers: a discount tier at $0.07 to $0.08, a mid-market at $0.75 to $1.26, an official-parity cluster at $1.39 to $1.40, and a premium tier at $2.10 charging for speed and platform integration.

Read this table with more suspicion than a normal price comparison, in both directions. A 30x gap between endpoints serving the same weights is not a stable market price; it is a market that has not converged. The discount tier may reflect promotional pricing, loss-leader routing, or quantized serving. The official-parity cluster at exactly $1.40 is eight providers pricing to the developer's list rather than to their own costs, which is its own kind of information: nobody in that cluster is competing on price at all.

What a real workload pays

A production assistant doing 10M input and 1M output tokens per day, no caching:

Provider tierDailyMonthlyAnnual
Official Z.AI ($1.40 / $4.40)$18.40$552$6,716
Cheapest host ($0.07 / $0.22)$0.92$27.60$336

Now the same workload with a 60 percent cache hit rate on input, typical for agents and RAG systems with stable prompts:

Provider tierDailyMonthlyAnnual
Official Z.AI ($0.26 cached)$11.56$346.80$4,219
Cheapest host ($0.013 cached)$0.58$17.34$211

On DeepSeek V3.2, cache terms flip the provider ranking, because the official API's cache discount beats third-party list prices. On GLM-5.2 they do not: the discount hosts offer cached input at $0.013 against the official $0.26, so caching widens the gap to 20x instead of closing it. Every model's provider decision has its own arithmetic, which is exactly why comparing list prices across a single dimension keeps picking wrong answers.

On this model, the entire decision reduces to one question: is the discount tier serving the product you think it is? If yes, the saving is enormous. If it is quantized, capped, or temporary, you have moved production onto a different product with a countdown timer. That question is unanswerable from a rate card, which is what blended benchmarks on the Mercatus Token Index and printed prices from actual trades are for.

GLM-5.2 vs DeepSeek V3.2 on cost

The two most widely hosted open-weight models, compared at their cheapest listed rates and official rates:

GLM-5.2DeepSeek V3.2
Official input / output $/1M$1.40 / $4.40$0.28 / $0.42
Cheapest listed input / output $/1M$0.07 / $0.22$0.20 / $0.32
Official cached input $/1M$0.26$0.028
Context window1M, no surcharge163K
Provider count20~10 major hosts
Cross-host input spread30x~3x

The comparison cuts both ways. At official rates, DeepSeek is 5x cheaper on input and 10x on output. At cheapest listed rates, GLM-5.2 is actually the cheaper model, with 6x the context window. Which one wins for a given workload depends on context needs, cache mix, and how much discount-tier risk the workload tolerates. What the table really shows is that "which model is cheaper" has no stable answer in a posted-price market; it changed twice in the last quarter and will change again.

The 95 percent collapse

The most important number on this page is not the spread, it is the velocity. GLM-5.2's cheapest available price fell from $1.40 to $0.07 per million input tokens in roughly 90 days, a 95 percent decline, while the official list price stayed exactly where it launched.

This is what open-weight hosting competition does to a model's price once enough providers pick it up, and it is the strongest argument yet for the pattern documented in Why Token Prices Differ: posted prices fragment because nothing forces them together. The practical consequences run in every direction. Anyone who locked an annual budget against launch pricing overpaid by an order of magnitude within a quarter. Anyone who committed to the cheapest host in May has been repriced twice since. And any provider holding the $1.40 line is now selling against competitors charging 5 percent of their rate.

The serving-cost floor underneath all of this is GPU economics: hardware cost per GPU-hour through realized throughput, the chain covered in Cost Per Token. Hosts running owned, high-utilization fleets can chase the floor; hosts on rented capacity cannot, which is a large part of why the tiers in the provider table look the way they do.

The first market-set forward price

Volatility like a 95 percent quarterly move is exactly why forward markets exist, and GLM-5.2 is where that mechanism reached AI tokens. In August 2026 it became the first model with a traded forward contract on the Mercatus Forward Market: price fixed at trade time, physically settled at expiry, meaning the buyer receives actual usable tokens rather than a cash difference.

For a model whose spot price moved 95 percent in a quarter, the value of the instrument does not need much explaining. A buyer can fix next month's GLM-5.2 cost today instead of guessing which pricing tier survives the next repricing. A provider can turn idle future capacity into signed revenue at a known rate. And every contract that trades prints a price, which begins building the thing this market has never had: a forward curve. The full mechanics, order book, settlement, and what the curve signals, are in How a Token Exchange Works.

How to choose a GLM-5.2 provider

  • Production workloads with quality sensitivity: official-parity tier, or verify the discount tier against your own evals before switching. At a 20x saving, the hour of benchmarking pays for itself immediately.
  • Batch and development workloads: discount tier, with rate limits checked first.
  • Long-context workloads: the 1M window with no surcharge is the model's standout economic feature, but confirm your host serves the full window; caps are common on third-party endpoints.
  • Heavy cache workloads: discount tier extends its lead here ($0.013 cached), the opposite of the DeepSeek dynamic.
  • Budget certainty: a forward contract fixes the price entirely, which no provider choice can do in a market that moved 95 percent last quarter.
  • Whatever you choose, recheck monthly. Twenty hosts and no convergence mechanism means rankings keep flipping.

Frequently asked questions

How much does GLM-5.2 cost per million tokens?

$1.40 input and $4.40 output on the official Z.AI API, with cached input at $0.26. Third-party hosts range from $0.07 to $2.10 input across 20 providers.

Who is the cheapest GLM-5.2 provider?

By list price, StreamLake at $0.07 input and $0.22 output, with OpenRouter routing at the same rate and Novita close behind. Verify precision, context limits, and rate limits before relying on discount-tier pricing for production.

Why is the GLM-5.2 price spread so large?

Twenty hosts, no price convergence mechanism, and a model young enough that the market has not settled. A 30x spread is what a token market looks like before transparency and arbitrage compress it; DeepSeek V3.2, older and more scrutinized, has already compressed to about 3x.

Has GLM-5.2 gotten cheaper?

The cheapest available price fell roughly 95 percent in the ~90 days after launch, from $1.40 to $0.07 per million input tokens, while the official rate held at $1.40.

Is GLM-5.2 cheaper than DeepSeek V3.2?

At official rates, no: DeepSeek is 5 to 10x cheaper. At cheapest listed rates, yes: $0.07 vs $0.20 input, with 6x the context window. The answer depends on your cache mix, context needs, and discount-tier risk tolerance.

What is a market-set forward price?

A price for future token delivery set by a trade on an exchange rather than by a provider's list price. GLM-5.2 is the first model with one: its forward contract on Mercatus is physically settled, delivering actual tokens at expiry.

Does GLM-5.2 charge more for long context?

No. The full 1M-token window bills at standard rates on the official API, with no long-context surcharge, unusual for models with windows this large.

Methodology

Provider pricing compiled from published rate cards and public pricing trackers (sourced from OpenRouter and Helicone data), August 2026. Benchmark percentiles from public aggregates. Tier groupings reflect listed prices, not verified serving precision; discount-tier endpoints should be independently verified before production use. Forward market facts reflect the Mercatus Forward Market's publicly announced first contracts, August 2026. Workload examples use stated token volumes and cache assumptions; recompute with your own mix. Last verified: 2026-08-10.