HomeBlogGLM-5.3 Flash API Pricing: $0.15 In, $0.50 Out per 1M Tokens
GeneralSep 21, 20269 min read

GLM-5.3 Flash API Pricing: $0.15 In, $0.50 Out per 1M Tokens

GLM-5.3 Flash costs $0.15 per million input tokens and $0.50 output on Z.ai's API, with cache hits at $0.03. The 50% launch promo ended September 9, but 11 of 30 hosts still sell below list and three still charge the promo rate. Every host, the workload math, and what a discount with an end date means for buyers.

M

Mercatus Compute

Author

GLM-5.3 Flash API Pricing: $0.15 In, $0.50 Out per 1M Tokens

GLM-5.3 Flash costs $0.15 per million input tokens, $0.50 per million output tokens, and $0.03 per million on cache hits on Z.ai’s official API as of September 21, 2026. That is the list price. For the first two weeks after launch on August 26, Z.ai sold it at half that, $0.075 input and $0.25 output, and the promotion ended on September 9.

The market did not end with it. Thirty providers serve the open weights on OpenRouter, and twelve days after Z.ai went back to list, eleven of them still price below it. Two, GMI Cloud and DeepInfra, are running the exact 50 percent promo Z.ai retired, labeled as a discount. A third, Open Inference, lists $0.075 and $0.25 with no label at all. This page covers the official rate card for the whole GLM line, every host’s price for GLM-5.3 Flash, the workload math, and why a price that comes with an end date is a different thing from a price.

GLM-5.3 Flash official API pricing

Z.ai’s current rate card for the GLM-5.3 generation. Prices in USD per million tokens as of September 21, 2026. Cached input storage is listed as free for a limited time on all four.

ModelInputCached inputOutput
GLM-5.3-Flash$0.15$0.03$0.50
GLM-5.3-FlashX$0.37$0.075$1.25
GLM-5.3$1.40$0.26$4.40
GLM-5.2$1.40$0.26$4.40

GLM-5.3 Flash is the small model of the generation: 320 billion total parameters with 18 billion active, a hybrid sparse and linear attention design, native image input, and a context window OpenRouter lists at 1.3M tokens. Z.ai published the weights under an MIT license on August 26, 2026, which is why the host table below exists. FlashX is the same model on a faster tier at 2.5x the price. GLM-5.3 and GLM-5.2, the flagships, sit at $1.40 and $4.40, nine times Flash on input and nearly nine times on output.

Two things on the rate card. The cache-hit rate is 20 percent of the input rate, a smaller discount than DeepSeek’s 2 percent cache rate on V4.1 Flash, so a cache-heavy agent workload gets less relief here. And there is no peak or off-peak billing on any Z.ai model; the rate holds around the clock. The comparison against DeepSeek, Qwen, and Kimi at the same tier is in the LLM API pricing comparison.

What the launch promo was, and where it went

From August 26 to September 9, Z.ai sold GLM-5.3 Flash at $0.075 input, $0.25 output, and $0.015 cached, half of list on every line. Any buyer who priced a workload on those numbers in the first two weeks saw the bill double on September 10 with no change to the model or the traffic.

That is the nature of a promotional rate, and it is worth naming. A list price is a number a seller can hold. A promotional price is a number a seller has already told you it will not hold. The date is part of the price.

What makes GLM-5.3 Flash a useful case is that the promo did not disappear when Z.ai ended it. It moved. The third-party market is still carrying it, in three forms:

Labeled discounts. GMI Cloud and DeepInfra list the full 50 percent off against a struck-through $0.15, so the buyer can see the end coming. Phala (15 percent), Novita (12 percent), and StreamLake (6 percent) are running smaller labeled discounts on the same list.

Unlabeled low prices. Open Inference at $0.075 and $0.25, Morph at $0.08 and $0.28, inference.net at $0.09 and $0.28, Relace at $0.099 and $0.33, and Wafer at $0.10 and $0.35 list below Z.ai with no promotion flag. Some of these may be permanent rates set by hosts with cheaper serving; some may be launch pricing that has not been marked as such. The listing does not say.

List price, from everyone else. Seventeen hosts, including Z.ai’s own endpoint, list exactly $0.15 and $0.50. Two are above it.

For a buyer, the practical question is which of the eleven below-list prices survives the next month. The labeled ones will end by definition. The unlabeled ones are a guess. The only price on the table that carries no expiration risk is the one Z.ai posts, and three hosts still charge half of it today.

GLM-5.3 Flash pricing by host

Prices per million tokens as listed on OpenRouter on September 21, 2026, sorted cheapest first. Hosts running a labeled discount show the discounted rate with the discount noted. Throughput and uptime are OpenRouter’s measurements on the same date.

HostInputOutputCache readThroughputUptimeNote
Open Inference$0.075$0.25$0.0210 tps98.70%
GMI Cloud$0.075$0.25$0.01523 tps98.61%50% off
DeepInfra$0.075$0.25$0.01514 tps99.02%50% off
Morph$0.08$0.28$0.01620 tps99.92%20% off
inference.net$0.09$0.28$0.0216 tps95.96%
Relace$0.099$0.33$0.019839 tps99.81%
Wafer$0.10$0.35$0.0221 tps99.86%
Phala$0.1275$0.425$0.025526 tps99.16%15% off
NovitaAI$0.132$0.44$0.026431 tps99.28%12% off
StreamLake$0.141$0.47$0.028219 tps99.11%6% off
Sail Research$0.1425$0.475$0.028524 tps99.39%
Baseten$0.15$0.50$0.0340 tps99.60%
Inceptron$0.15$0.50$0.0346 tps98.51%
NEAR AI$0.15$0.50$0.03524 tps98.05%
CoreWeave$0.15$0.50$0.0561 tps99.75%
AtlasCloud$0.15$0.50$0.0326 tps99.20%
Fireworks$0.15$0.50$0.0354 tps99.60%
Friendli$0.15$0.50$0.0321 tps99.41%
SiliconFlow$0.15$0.50$0.0330 tps99.51%
DigitalOcean$0.15$0.50$0.0331 tps95.85%
Together$0.15$0.50$0.0368 tps99.25%
Reka AI$0.15$0.50$0.0310 tps99.56%
Parasail$0.15$0.50$0.0347 tps98.95%
Venice$0.15$0.50$0.0316 tps98.20%
io.net$0.15$0.50$0.0316 tps99.45%
Cloudflare$0.15$0.50$0.0341 tps99.75%
Z.ai (official)$0.15$0.50$0.0336 tps99.58%
Crusoe$0.15$0.50$0.0361 tps93.81%
NextBit$0.18$0.60$0.03640 tps97.71%
Modal$0.45$1.50$0.0940 tps99.78%

The spread is 6x on input ($0.075 to $0.45) and 6x on output ($0.25 to $1.50). Take Modal out and it is 2.4x. The shape is different from DeepSeek’s market, where hosts scattered across the whole range after the August repricing, and from Qwen’s, where six of seven hosts sat at one number. Here the market has a hard center: seventeen hosts at exactly Z.ai’s list, with a tail of eleven below it that is mostly promotional.

Throughput is the other axis, and it does not track price. Together at 68 tokens per second and CoreWeave and Crusoe at 61 all charge list. The three cheapest hosts run at 10 to 23 tokens per second. A buyer paying half price at Open Inference is getting a seventh of the throughput of a list-price host, which is the trade the discount is buying. Why hosts of the same weights price so differently is covered in Why token prices differ.

What a real workload pays

A production assistant doing 10 million input tokens and 1 million output tokens a day, uncached, at list rates on September 21, 2026. GLM-5.3 Flash is compared against the tier above it, the promo rate it launched at, and the two closest open-weight small models.

RouteDailyAnnual
GLM-5.3 Flash, launch promo (ended Sept 9)$1.00$365
GLM-5.3 Flash, GMI Cloud or DeepInfra (50% off, current)$1.00$365
GLM-5.3 Flash, Morph (20% off)$1.08$394
GLM-5.3 Flash, inference.net$1.18$431
GLM-5.3 Flash, Z.ai official, 60% cache hits$1.28$467
GLM-5.3 Flash, Z.ai official, uncached$2.00$730
DeepSeek V4.1 Flash, official off-peak$2.10$767
Qwen3.8 Flash, official$1.97$719
GLM-5.3 Flash, Modal$6.00$2,190
GLM-5.3 FlashX, Z.ai$4.95$1,807
GLM-5.3 (flagship), Z.ai$18.40$6,716

Three things fall out of the table.

At list, the three open-weight small models are within 7 percent of each other. GLM-5.3 Flash at $2.00 a day, Qwen3.8 Flash at $1.97, DeepSeek V4.1 Flash at $2.10 off-peak. Three developers, three architectures, and the official price of a million tokens landed in the same place. That is what a competitive tier looks like, and it means the choice between them is a quality and throughput decision, not a price one. The DeepSeek numbers are in DeepSeek V4.1 Flash pricing and the Qwen numbers in Qwen API pricing.

The promo was worth $365 a year on this workload, and it is still available, for now. A buyer who moved to GMI Cloud or DeepInfra when Z.ai ended its promotion kept the $1.00 a day. When those hosts end theirs, the same buyer is back to $2.00 unless another host picks it up. That is a workload whose annual cost is set by a sequence of expiration dates the buyer does not control.

The flagship is nine times the price of Flash. At $18.40 a day, GLM-5.3 costs more than nine Flash workloads. Z.ai’s own positioning, and the FlashX tier at $4.95 in between, says the company expects most volume to sit at the bottom of that ladder.

Which GLM model and host to use

High-volume, routine tokens: GLM-5.3 Flash from a list-price host with the throughput you need. Together, CoreWeave, and Fireworks serve it at 54 to 68 tokens per second at $0.15 and $0.50. The discount hosts are cheaper and slower; use them for batch work that does not care about latency, and re-check the price each month.

Locking a budget: price it at list, not at the promo. $2.00 a day on the 10M/1M workload is the number that holds. Anything below it is a discount that ends, and the Mercatus Token Index tracks where the traded price actually settles as the promotional tail rolls off.

Cache-heavy agents: Z.ai’s $0.03 cache rate is matched by most hosts, and CoreWeave’s $0.05 and NEAR AI’s $0.035 are the ones to avoid if cache hits are most of your input. At 60 percent cache hits the official route drops to $1.28 a day.

Need the flagship: GLM-5.3 at $1.40 and $4.40, the same price as GLM-5.2, which still sells alongside it. The GLM-5.2 host market and its launch-month spread are in GLM-5.2 API pricing.

Latency over everything: Modal at 0.29 seconds first-token latency, at 3x list. That is a premium for a specific need, and it is the only host on the table charging one.

Why a promotional price is not a price

GLM-5.3 Flash launched with the clearest example this year of a rate that was never meant to last. Z.ai said so on day one: 50 percent off through September 9. Buyers who read the date knew what their cost would be on September 10. Buyers who read the number did not.

The third-party market has now inherited that structure. Two hosts are running the same promotion with the same label and no stated end date. Three more are running smaller ones. Five list below Z.ai with no label, which is either a permanent price or a promotion nobody has flagged. Every one of those eleven prices is a number the seller can change by announcement, and the only signal a buyer gets is the day it changes.

That is the gap a forward market closes. A buyer who needs GLM-5.3 Flash capacity for the next twelve months at a known cost cannot get that from a rate card, promotional or not. They can get it from a contract that locks the capacity before the discount rolls off. How to lock in token prices covers what that looks like, and How a token exchange works covers how a price set by trades, rather than by a seller’s calendar, gets made.

Frequently asked questions

How much does GLM-5.3 Flash cost? As of September 21, 2026, Z.ai’s official API charges $0.15 per million input tokens, $0.50 per million output tokens, and $0.03 per million on cache hits. Third-party hosts on OpenRouter list it from $0.075 input and $0.25 output (Open Inference, GMI Cloud, DeepInfra) up to $0.45 and $1.50 (Modal).

Did the GLM-5.3 Flash launch discount end? On Z.ai’s own API, yes, on September 9, 2026. The launch rate was $0.075 input and $0.25 output. As of September 21, GMI Cloud and DeepInfra still list that rate as a 50 percent discount, and Open Inference lists it without a label.

Is GLM-5.3 Flash open-weight? Yes. Z.ai released the weights under an MIT license on August 26, 2026. Thirty providers serve them on OpenRouter as of September 21.

What is the difference between GLM-5.3 Flash and GLM-5.3 FlashX? Same model, faster serving tier. FlashX costs $0.37 input and $1.25 output on Z.ai, about 2.5x the standard Flash rate.

How does GLM-5.3 Flash pricing compare to DeepSeek V4.1 Flash? At list, they are close. GLM-5.3 Flash is $0.15 and $0.50; DeepSeek V4.1 Flash is $0.15 and $0.60 off-peak and $0.30 and $1.20 at peak. GLM has no peak pricing. DeepSeek’s cache-hit rate is lower ($0.003 versus $0.03), which favors DeepSeek on cache-heavy agent workloads.

Who is the cheapest GLM-5.3 Flash host? On September 21, 2026: Open Inference, GMI Cloud, and DeepInfra at $0.075 input and $0.25 output. Two of the three are labeled promotions. Among hosts at or near list, Together and CoreWeave have the highest throughput at 68 and 61 tokens per second.

Does Z.ai have peak and off-peak pricing? No. All Z.ai rates are flat around the clock. Among the open-weight models tracked on Mercatus, DeepSeek is the only one that bills by time of day.

Methodology

Official GLM-5.3 Flash, FlashX, GLM-5.3, and GLM-5.2 rates are from Z.ai’s developer pricing page on September 21, 2026. The launch promotional rate and its September 9 end date are from Z.ai’s launch announcement as reported at the time and from the promotional labels still visible on third-party listings. Third-party host prices, discount labels, throughput, latency, and uptime are from OpenRouter’s GLM 5.3 Flash provider table on September 21, 2026; two providers with duplicate endpoints are listed once. Model specifications are from Z.ai’s release materials and OpenRouter’s listing. Workload calculations assume 10 million input and 1 million output tokens per day, uncached unless stated, times 365; DeepSeek and Qwen comparison figures are from the corresponding Mercatus pricing pages. Prices change; labeled discounts in particular should be expected to end.