HomeBlogQwen API Pricing: Every Qwen3.8 Model and Host
GeneralSep 7, 20269 min read

Qwen API Pricing: Every Qwen3.8 Model and Host

Qwen3.8 Max costs $2 input and $6 output per million tokens. Its open-weight twin costs the same at every host, while Qwen3.8 27B spans 3x across 13 hosts.

M

Mercatus Compute

Author

Qwen API Pricing: Every Qwen3.8 Model and Host

Qwen3.8 Max, Alibaba's flagship, costs $2.00 per million input tokens and $6.00 per million output tokens, with cached input at $0.25. Qwen3.8 Flash costs $0.15 and $0.47. Qwen3.8 27B, the open-weight dense model, lists at $0.15 and $2.00 from the cheapest host and up to $0.45 and $3.20 from the most expensive. And the open-weight version of the flagship, Qwen3.8 2.4T A95B, costs exactly $2.00 and $6.00 from six of the seven hosts that serve it, which makes it the first major open-weight model whose third-party market has no spread at all.

That last fact is the reason Qwen deserves its own page rather than a row in a comparison table. DeepSeek's open weights trade at an 8x spread across hosts. GLM-5.2's traded at 30x in its launch month. Qwen's flagship open weights trade at 1x. Same market structure, many sellers of identical weights, and a completely different price outcome. This page covers the official rates for the whole Qwen line, the host tables for the two open-weight 3.8 models, the workload math, and why the spread looks the way it does.

Qwen3.8 official API pricing

Alibaba sells Qwen through Model Studio, with international-region rates listed below. All prices are USD per million tokens as of September 7, 2026.

ModelInputCached inputOutputContextWeightsReleased
Qwen3.8 Max (0902)$2.00$0.25$6.001MClosedSep 3, 2026
Qwen3.8 2.4T A95B$2.00$0.25$6.001MOpenAug 12, 2026
Qwen3.8 Flash$0.15$0.016$0.471MClosedAug 26, 2026
Qwen3.8 27B$0.425$0.085$2.551MOpenAug 14, 2026

Qwen3.8 Max is a 2.4-trillion-parameter mixture-of-experts model with about 95 billion active parameters per token, a 1M-token context window, and text, image, and video input. It went generally available on August 3, 2026 at $2.00 and $6.00, replacing a credit-only subscription plan with per-token pricing, and the 0902 build shipped on September 3. The 2.4T A95B open-weight release on August 12 is the same architecture with published weights. Qwen3.8 27B is a dense open-weight vision-language model. Flash is the small closed model.

Three things to note on the official rate card. Cached input on Max is $0.25, an 87.5 percent discount, less steep than DeepSeek's 97 percent but meaningful on agent workloads. There is no peak and off-peak billing on any Qwen model; the rates hold around the clock, which matters now that DeepSeek bills by time of day. And on the 27B, Alibaba's own rate of $0.425 input is near the top of the third-party range, not the bottom.

The full Qwen lineup, by generation

Qwen ships a new generation roughly every two to three months, and older generations stay on the API at their old prices. That leaves a long menu. Prices below are the base listings on OpenRouter as of September 7, 2026, which for the closed models is Alibaba's own rate.

ModelInputOutputContextReleased
Qwen3.8 Max$2.00$6.001MSep 2026
Qwen3.8 2.4T A95B (open)$2.00$6.001MAug 2026
Qwen3.8 27B (open)$0.15$2.001MAug 2026
Qwen3.8 Flash$0.15$0.471MAug 2026
Qwen3.7 Max$1.475$4.4251MMay 2026
Qwen3.7 Plus$0.32$1.281MJun 2026
Qwen3.7 Flash$0.03$0.131MJul 2026
Qwen3.6 Plus$0.325$1.951MApr 2026
Qwen3.6 27B (open)$0.30$2.00262KApr 2026
Qwen3.6 35B A3B (open)$0.05$0.70262KApr 2026
Qwen3.5 397B A17B (open)$0.39$2.34262KFeb 2026
Qwen3.5 Plus$0.30$1.801MApr 2026
Qwen3.5 Flash$0.065$0.261MFeb 2026
Qwen3 Max$0.78$3.90262KSep 2025
Qwen3 Coder Plus$0.65$3.251MSep 2025
Qwen3 235B A22B Instruct (open)$0.0875$0.35262KJul 2025

The price ladder inside Qwen alone runs from $0.03 input on Qwen3.7 Flash to $2.00 on the 3.8 flagship, a 67x range. Each generation's flagship has also repriced upward: Qwen3 Max at $0.78, Qwen3.7 Max at $1.475, Qwen3.8 Max at $2.00. That is the opposite of the DeepSeek pattern through 2025, where each release cut the price, and worth remembering when a "Qwen is cheap" assumption is ported from one generation to the next.

Qwen3.8 2.4T A95B: six hosts, one price

Prices as listed on OpenRouter on September 7, 2026.

HostInputOutputCache readThroughputUptime
Together$2.00$6.00$0.25103 tps100%
Modal$2.00$6.00$0.25127 tps99.96%
DeepInfra$2.00$6.00$0.2083 tps97.97%
SiliconFlow$2.00$6.00$0.2530 tps99.98%
Alibaba Cloud (official)$2.00$6.00$0.2539 tps99.97%
Novita$2.00$6.00$0.2535 tps99.99%
Venice$2.50$7.50$0.312587 tps100%

Six of seven hosts list the identical price to the cent. The only variation is Venice at a 25 percent premium and DeepInfra's slightly cheaper cache read. On the same day, four of these hosts (DeepInfra, SiliconFlow, Novita, Alibaba Cloud) also served DeepSeek V4 Pro at four different prices spanning a 23 percent range on input. The weights are open, the hosts compete on everything else, and nobody has cut the price.

Two readings are possible. One is serving cost: 95 billion active parameters per token, a 1M window, and multimodal input make this an expensive model to run, and $2.00 and $6.00 may be close to what it costs a third party at reasonable utilization. The other is anchoring: the open weights shipped nine days after the closed flagship at the flagship's price, and in the first month hosts matched the list rather than undercutting it. If it is anchoring, the spread opens as hosts optimize. If it is cost, it stays tight. The DeepSeek V4 Pro market took about three weeks to re-sort after its August repricing, so the next month of this table is the tell.

What already differs is throughput. Modal at 127 tokens per second and Together at 103 serve the same tokens, at the same price, roughly three times faster than Alibaba's own endpoint at 39. On this model the routing decision is entirely about speed and uptime, not price.

Qwen3.8 27B by host

The dense 27B model is where Qwen looks like a normal open-weight market. Thirteen hosts, sorted by input price, as listed on OpenRouter on September 7, 2026.

HostInputOutputCache readThroughputUptime
Darkbloom$0.15$2.00n/a32 tps99.80%
Reka AI$0.214$2.55$0.1544 tps97.09%
Parasail$0.24$2.20$0.0578 tps99.27%
AkashML$0.25$2.20$0.0559 tps98.98%
Phala$0.30$3.00$0.0549 tps98.29%
Chutes$0.32$2.50$0.03226 tps98.46%
Ionstream$0.35$2.55$0.0553 tps98.01%
CoreWeave$0.40$3.00$0.1545 tps99.97%
Novita$0.42$3.00$0.08535 tps99.82%
Alibaba Cloud (official)$0.425$2.55$0.08553 tps99.98%
io.net$0.432$3.06$0.22548 tps99.92%
Venice$0.45$3.20n/a81 tps98.82%
Cloudflare$0.45$3.20$0.0536 tps87.94%

The spread is 3x on input ($0.15 to $0.45) and 1.6x on output ($2.00 to $3.20). Alibaba's official API sits tenth of thirteen on input. This is the GLM pattern rather than the DeepSeek pattern: the developer prices its own API at a premium and lets third parties fight for volume, which is the structure covered in Why token prices differ.

Note the output-heavy shape of this model's pricing. At $2.00 output against $0.15 input, the 27B's output tokens cost 13x its input tokens, versus 3x on Qwen3.8 Max and Flash. A generation-heavy workload on the 27B can cost more than the same workload on Flash, as the table below shows.

What a real workload pays

A production assistant doing 10M input and 1M output tokens per day, uncached, at the cheapest listed rate for each model. DeepSeek, GLM, and Kimi rows are from their Mercatus pricing pages for comparison.

Model and routeDailyAnnual
Qwen3.7 Flash$0.43$157
DeepSeek V4 Flash (cheapest host)$0.85$309
Qwen3.8 Flash$1.97$719
DeepSeek V4 Flash (official, off-peak)$2.86$1,044
Qwen3.8 27B (Darkbloom)$3.50$1,278
Qwen3.7 Plus$4.48$1,635
GLM-5.2 (cheapest host)$5.28$1,927
Qwen3.8 27B (Alibaba official)$6.80$2,482
DeepSeek V4 Pro (official, off-peak)$8.58$3,132
GLM-5.2 (official)$18.40$6,716
Qwen3.7 Max$19.18$6,999
Qwen3.8 Max or 2.4T A95B (any host)$26.00$9,490
Kimi K3$45.00$16,425

Three things fall out of the table.

Qwen3.8 Flash is the value play in the 3.8 line, not the 27B. On an input-heavy workload Flash costs $1.97 a day against $3.50 for the cheapest 27B host, because the 27B's $2.00 output rate is 4x Flash's. The 27B earns its price only when the task needs open weights, a dense architecture, or vision.

Qwen3.8 Max is a premium model priced like one. At $26 a day it costs 3x DeepSeek V4 Pro off-peak and 1.4x GLM-5.2 official. Caching helps: at a 60 percent cache hit rate the same workload drops to $15.50 a day, a 40 percent cut. It does not change the tier.

The cheapest Qwen is a generation old. Qwen3.7 Flash at $0.03 and $0.13 runs the workload for $0.43 a day, half the price of the cheapest DeepSeek V4 Flash host and a fifth of Qwen3.8 Flash. For classification, extraction, and routing tasks where the 3.8 quality gain does not show up in evals, the 3.7 Flash rate is the number to beat.

Which Qwen model and host to use

Flagship reasoning, coding, or multimodal work: Qwen3.8 Max or the 2.4T open weights, same price either way. Pick the host on throughput. Modal and Together serve it 3x faster than Alibaba at the same rate. Use the open weights if you need a second source or self-hosting as a fallback.

Open weights on a budget: Qwen3.8 27B from Darkbloom, Parasail, or AkashML at $0.15 to $0.25 input. Check uptime and throughput before committing; the cheapest host on this list is also one of the slowest.

High-volume, low-stakes tokens: Qwen3.8 Flash at $0.15 and $0.47 if you need the 3.8 generation, Qwen3.7 Flash at $0.03 and $0.13 if you do not. Run the eval before paying the 5x difference.

Cache-heavy agents: Max's $0.25 cached rate and Flash's $0.016 are both official-API terms; among 27B hosts, Chutes ($0.032) and the $0.05 tier (Parasail, AkashML, Phala, Ionstream, Cloudflare) beat Alibaba's $0.085.

The model-by-model comparison against DeepSeek, GLM, and Kimi is in the LLM API pricing comparison.

Why Qwen's spread looks nothing like DeepSeek's

Three open-weight Chinese flagships, three market structures. DeepSeek priced its own API at the floor for months, then raised it, and the third-party market split around the move. GLM-5.2 launched with a $0.07 promotional tier under a $1.40 official rate and watched the spread collapse within two weeks. Qwen3.8's open flagship launched at the closed flagship's price and every host but one held it.

The difference is not the weights being open. In all three cases they are. It is the size of the model relative to what hosts can serve profitably, and the price the developer set before the third parties arrived. A 2.4T-parameter model with 95B active per token leaves little room under $2.00 and $6.00; a 27B dense model leaves plenty, which is why the 27B market spread to 3x in three weeks while the 2.4T market stayed at 1x.

For buyers, the practical point is the same one every model page on this site ends up making: the official rate card tells you what one seller wants, not what the model costs. On Qwen3.8 Max there is no cheaper seller today. On Qwen3.8 27B the official seller is one of the more expensive ones. Neither fact is visible from Alibaba's pricing page, and both can change by the month. The Mercatus Token Index tracks where these models actually trade, and How a token exchange works covers what it takes to turn eighteen rate cards into one price.

Frequently asked questions

How much does the Qwen API cost?
As of September 2026, Qwen3.8 Max costs $2.00 per million input tokens and $6.00 per million output tokens on Alibaba Cloud Model Studio, with cached input at $0.25. Qwen3.8 Flash costs $0.15 and $0.47. The open-weight Qwen3.8 27B ranges from $0.15 to $0.45 input and $2.00 to $3.20 output across 13 third-party hosts. Older generations stay available: Qwen3.7 Max at $1.475 and $4.425, Qwen3.7 Flash at $0.03 and $0.13.

Is Qwen3.8 Max open-weight?
The Max endpoint is proprietary, but Alibaba released Qwen3.8 2.4T A95B on August 12, 2026 as the open-weight variant of the same architecture. Seven hosts serve it, six of them at the same $2.00 and $6.00 as the official Max endpoint.

What is the cheapest Qwen model?
Qwen3.7 Flash at $0.03 input and $0.13 output per million tokens. In the current generation, Qwen3.8 Flash at $0.15 and $0.47.

Who is the cheapest host for Qwen3.8 27B?
Darkbloom at $0.15 input and $2.00 output on September 7, 2026, followed by Reka AI, Parasail, and AkashML. Alibaba's official API at $0.425 is tenth of thirteen on input.

Is Qwen cheaper than DeepSeek?
Not at the flagship tier. Qwen3.8 Max costs $2.00 and $6.00 against DeepSeek V4 Pro's $0.66 and $1.98 off-peak, about 3x more. At the small-model tier they are close: Qwen3.8 Flash at $0.15 and $0.47 against DeepSeek V4 Flash at $0.22 and $0.66 official, or $0.0679 and $0.168 from the cheapest DeepSeek host. Qwen3.7 Flash at $0.03 and $0.13 undercuts everything.

Does Qwen have peak and off-peak pricing?
No. As of September 2026 all Qwen rates are flat around the clock. Among the models tracked on Mercatus, DeepSeek is the only one billing by time of day.

Why do all the hosts charge the same for Qwen3.8 2.4T?
The open weights launched nine days after the closed flagship at the flagship's exact price, and in the first month no host has undercut it. Whether that reflects serving cost for a 2.4T-parameter model or launch-window anchoring will show in whether the spread opens over the next month.

Methodology

Official Qwen rates are Alibaba Cloud Model Studio international-region list prices as reflected on OpenRouter's Alibaba Cloud International provider listings on September 7, 2026, including cache-read rates; China-region pricing may differ and was not reviewed. Third-party host prices, throughput, and uptime for Qwen3.8 2.4T A95B and Qwen3.8 27B are from OpenRouter's model pages on the same date and reflect list prices including any promotions in effect. Prices for older Qwen generations are OpenRouter base listings on the same date. Release dates are from OpenRouter listing dates and Alibaba's Qwen3.8 Max availability announcement of August 3, 2026. Workload calculations assume uncached input; the cached scenario assumes a 60 percent cache hit rate on input. DeepSeek, GLM-5.2, and Kimi K3 comparison figures are from the corresponding Mercatus pricing pages.