Qwen3.8 Max, Alibaba's flagship, costs $2.00 per million input tokens and $6.00 per million output tokens, with cached input at $0.25. Qwen3.8 Flash costs $0.15 and $0.47. Qwen3.8 27B, the open-weight dense model, lists at $0.15 and $2.00 from the cheapest host and up to $0.45 and $3.20 from the most expensive. And the open-weight version of the flagship, Qwen3.8 2.4T A95B, costs exactly $2.00 and $6.00 from six of the seven hosts that serve it, which makes it the first major open-weight model whose third-party market has no spread at all.
That last fact is the reason Qwen deserves its own page rather than a row in a comparison table. DeepSeek's open weights trade at an 8x spread across hosts. GLM-5.2's traded at 30x in its launch month. Qwen's flagship open weights trade at 1x. Same market structure, many sellers of identical weights, and a completely different price outcome. This page covers the official rates for the whole Qwen line, the host tables for the two open-weight 3.8 models, the workload math, and why the spread looks the way it does.
Qwen3.8 official API pricing
Alibaba sells Qwen through Model Studio, with international-region rates listed below. All prices are USD per million tokens as of September 7, 2026.
| Model | Input | Cached input | Output | Context | Weights | Released |
|---|---|---|---|---|---|---|
| Qwen3.8 Max (0902) | $2.00 | $0.25 | $6.00 | 1M | Closed | Sep 3, 2026 |
| Qwen3.8 2.4T A95B | $2.00 | $0.25 | $6.00 | 1M | Open | Aug 12, 2026 |
| Qwen3.8 Flash | $0.15 | $0.016 | $0.47 | 1M | Closed | Aug 26, 2026 |
| Qwen3.8 27B | $0.425 | $0.085 | $2.55 | 1M | Open | Aug 14, 2026 |
Qwen3.8 Max is a 2.4-trillion-parameter mixture-of-experts model with about 95 billion active parameters per token, a 1M-token context window, and text, image, and video input. It went generally available on August 3, 2026 at $2.00 and $6.00, replacing a credit-only subscription plan with per-token pricing, and the 0902 build shipped on September 3. The 2.4T A95B open-weight release on August 12 is the same architecture with published weights. Qwen3.8 27B is a dense open-weight vision-language model. Flash is the small closed model.
Three things to note on the official rate card. Cached input on Max is $0.25, an 87.5 percent discount, less steep than DeepSeek's 97 percent but meaningful on agent workloads. There is no peak and off-peak billing on any Qwen model; the rates hold around the clock, which matters now that DeepSeek bills by time of day. And on the 27B, Alibaba's own rate of $0.425 input is near the top of the third-party range, not the bottom.
The full Qwen lineup, by generation
Qwen ships a new generation roughly every two to three months, and older generations stay on the API at their old prices. That leaves a long menu. Prices below are the base listings on OpenRouter as of September 7, 2026, which for the closed models is Alibaba's own rate.
| Model | Input | Output | Context | Released |
|---|---|---|---|---|
| Qwen3.8 Max | $2.00 | $6.00 | 1M | Sep 2026 |
| Qwen3.8 2.4T A95B (open) | $2.00 | $6.00 | 1M | Aug 2026 |
| Qwen3.8 27B (open) | $0.15 | $2.00 | 1M | Aug 2026 |
| Qwen3.8 Flash | $0.15 | $0.47 | 1M | Aug 2026 |
| Qwen3.7 Max | $1.475 | $4.425 | 1M | May 2026 |
| Qwen3.7 Plus | $0.32 | $1.28 | 1M | Jun 2026 |
| Qwen3.7 Flash | $0.03 | $0.13 | 1M | Jul 2026 |
| Qwen3.6 Plus | $0.325 | $1.95 | 1M | Apr 2026 |
| Qwen3.6 27B (open) | $0.30 | $2.00 | 262K | Apr 2026 |
| Qwen3.6 35B A3B (open) | $0.05 | $0.70 | 262K | Apr 2026 |
| Qwen3.5 397B A17B (open) | $0.39 | $2.34 | 262K | Feb 2026 |
| Qwen3.5 Plus | $0.30 | $1.80 | 1M | Apr 2026 |
| Qwen3.5 Flash | $0.065 | $0.26 | 1M | Feb 2026 |
| Qwen3 Max | $0.78 | $3.90 | 262K | Sep 2025 |
| Qwen3 Coder Plus | $0.65 | $3.25 | 1M | Sep 2025 |
| Qwen3 235B A22B Instruct (open) | $0.0875 | $0.35 | 262K | Jul 2025 |
The price ladder inside Qwen alone runs from $0.03 input on Qwen3.7 Flash to $2.00 on the 3.8 flagship, a 67x range. Each generation's flagship has also repriced upward: Qwen3 Max at $0.78, Qwen3.7 Max at $1.475, Qwen3.8 Max at $2.00. That is the opposite of the DeepSeek pattern through 2025, where each release cut the price, and worth remembering when a "Qwen is cheap" assumption is ported from one generation to the next.
Qwen3.8 2.4T A95B: six hosts, one price
Prices as listed on OpenRouter on September 7, 2026.
| Host | Input | Output | Cache read | Throughput | Uptime |
|---|---|---|---|---|---|
| Together | $2.00 | $6.00 | $0.25 | 103 tps | 100% |
| Modal | $2.00 | $6.00 | $0.25 | 127 tps | 99.96% |
| DeepInfra | $2.00 | $6.00 | $0.20 | 83 tps | 97.97% |
| SiliconFlow | $2.00 | $6.00 | $0.25 | 30 tps | 99.98% |
| Alibaba Cloud (official) | $2.00 | $6.00 | $0.25 | 39 tps | 99.97% |
| Novita | $2.00 | $6.00 | $0.25 | 35 tps | 99.99% |
| Venice | $2.50 | $7.50 | $0.3125 | 87 tps | 100% |
Six of seven hosts list the identical price to the cent. The only variation is Venice at a 25 percent premium and DeepInfra's slightly cheaper cache read. On the same day, four of these hosts (DeepInfra, SiliconFlow, Novita, Alibaba Cloud) also served DeepSeek V4 Pro at four different prices spanning a 23 percent range on input. The weights are open, the hosts compete on everything else, and nobody has cut the price.
Two readings are possible. One is serving cost: 95 billion active parameters per token, a 1M window, and multimodal input make this an expensive model to run, and $2.00 and $6.00 may be close to what it costs a third party at reasonable utilization. The other is anchoring: the open weights shipped nine days after the closed flagship at the flagship's price, and in the first month hosts matched the list rather than undercutting it. If it is anchoring, the spread opens as hosts optimize. If it is cost, it stays tight. The DeepSeek V4 Pro market took about three weeks to re-sort after its August repricing, so the next month of this table is the tell.
What already differs is throughput. Modal at 127 tokens per second and Together at 103 serve the same tokens, at the same price, roughly three times faster than Alibaba's own endpoint at 39. On this model the routing decision is entirely about speed and uptime, not price.
Qwen3.8 27B by host
The dense 27B model is where Qwen looks like a normal open-weight market. Thirteen hosts, sorted by input price, as listed on OpenRouter on September 7, 2026.
| Host | Input | Output | Cache read | Throughput | Uptime |
|---|---|---|---|---|---|
| Darkbloom | $0.15 | $2.00 | n/a | 32 tps | 99.80% |
| Reka AI | $0.214 | $2.55 | $0.15 | 44 tps | 97.09% |
| Parasail | $0.24 | $2.20 | $0.05 | 78 tps | 99.27% |
| AkashML | $0.25 | $2.20 | $0.05 | 59 tps | 98.98% |
| Phala | $0.30 | $3.00 | $0.05 | 49 tps | 98.29% |
| Chutes | $0.32 | $2.50 | $0.032 | 26 tps | 98.46% |
| Ionstream | $0.35 | $2.55 | $0.05 | 53 tps | 98.01% |
| CoreWeave | $0.40 | $3.00 | $0.15 | 45 tps | 99.97% |
| Novita | $0.42 | $3.00 | $0.085 | 35 tps | 99.82% |
| Alibaba Cloud (official) | $0.425 | $2.55 | $0.085 | 53 tps | 99.98% |
| io.net | $0.432 | $3.06 | $0.225 | 48 tps | 99.92% |
| Venice | $0.45 | $3.20 | n/a | 81 tps | 98.82% |
| Cloudflare | $0.45 | $3.20 | $0.05 | 36 tps | 87.94% |
The spread is 3x on input ($0.15 to $0.45) and 1.6x on output ($2.00 to $3.20). Alibaba's official API sits tenth of thirteen on input. This is the GLM pattern rather than the DeepSeek pattern: the developer prices its own API at a premium and lets third parties fight for volume, which is the structure covered in Why token prices differ.
Note the output-heavy shape of this model's pricing. At $2.00 output against $0.15 input, the 27B's output tokens cost 13x its input tokens, versus 3x on Qwen3.8 Max and Flash. A generation-heavy workload on the 27B can cost more than the same workload on Flash, as the table below shows.
What a real workload pays
A production assistant doing 10M input and 1M output tokens per day, uncached, at the cheapest listed rate for each model. DeepSeek, GLM, and Kimi rows are from their Mercatus pricing pages for comparison.
| Model and route | Daily | Annual |
|---|---|---|
| Qwen3.7 Flash | $0.43 | $157 |
| DeepSeek V4 Flash (cheapest host) | $0.85 | $309 |
| Qwen3.8 Flash | $1.97 | $719 |
| DeepSeek V4 Flash (official, off-peak) | $2.86 | $1,044 |
| Qwen3.8 27B (Darkbloom) | $3.50 | $1,278 |
| Qwen3.7 Plus | $4.48 | $1,635 |
| GLM-5.2 (cheapest host) | $5.28 | $1,927 |
| Qwen3.8 27B (Alibaba official) | $6.80 | $2,482 |
| DeepSeek V4 Pro (official, off-peak) | $8.58 | $3,132 |
| GLM-5.2 (official) | $18.40 | $6,716 |
| Qwen3.7 Max | $19.18 | $6,999 |
| Qwen3.8 Max or 2.4T A95B (any host) | $26.00 | $9,490 |
| Kimi K3 | $45.00 | $16,425 |
Three things fall out of the table.
Qwen3.8 Flash is the value play in the 3.8 line, not the 27B. On an input-heavy workload Flash costs $1.97 a day against $3.50 for the cheapest 27B host, because the 27B's $2.00 output rate is 4x Flash's. The 27B earns its price only when the task needs open weights, a dense architecture, or vision.
Qwen3.8 Max is a premium model priced like one. At $26 a day it costs 3x DeepSeek V4 Pro off-peak and 1.4x GLM-5.2 official. Caching helps: at a 60 percent cache hit rate the same workload drops to $15.50 a day, a 40 percent cut. It does not change the tier.
The cheapest Qwen is a generation old. Qwen3.7 Flash at $0.03 and $0.13 runs the workload for $0.43 a day, half the price of the cheapest DeepSeek V4 Flash host and a fifth of Qwen3.8 Flash. For classification, extraction, and routing tasks where the 3.8 quality gain does not show up in evals, the 3.7 Flash rate is the number to beat.
Which Qwen model and host to use
Flagship reasoning, coding, or multimodal work: Qwen3.8 Max or the 2.4T open weights, same price either way. Pick the host on throughput. Modal and Together serve it 3x faster than Alibaba at the same rate. Use the open weights if you need a second source or self-hosting as a fallback.
Open weights on a budget: Qwen3.8 27B from Darkbloom, Parasail, or AkashML at $0.15 to $0.25 input. Check uptime and throughput before committing; the cheapest host on this list is also one of the slowest.
High-volume, low-stakes tokens: Qwen3.8 Flash at $0.15 and $0.47 if you need the 3.8 generation, Qwen3.7 Flash at $0.03 and $0.13 if you do not. Run the eval before paying the 5x difference.
Cache-heavy agents: Max's $0.25 cached rate and Flash's $0.016 are both official-API terms; among 27B hosts, Chutes ($0.032) and the $0.05 tier (Parasail, AkashML, Phala, Ionstream, Cloudflare) beat Alibaba's $0.085.
The model-by-model comparison against DeepSeek, GLM, and Kimi is in the LLM API pricing comparison.
Why Qwen's spread looks nothing like DeepSeek's
Three open-weight Chinese flagships, three market structures. DeepSeek priced its own API at the floor for months, then raised it, and the third-party market split around the move. GLM-5.2 launched with a $0.07 promotional tier under a $1.40 official rate and watched the spread collapse within two weeks. Qwen3.8's open flagship launched at the closed flagship's price and every host but one held it.
The difference is not the weights being open. In all three cases they are. It is the size of the model relative to what hosts can serve profitably, and the price the developer set before the third parties arrived. A 2.4T-parameter model with 95B active per token leaves little room under $2.00 and $6.00; a 27B dense model leaves plenty, which is why the 27B market spread to 3x in three weeks while the 2.4T market stayed at 1x.
For buyers, the practical point is the same one every model page on this site ends up making: the official rate card tells you what one seller wants, not what the model costs. On Qwen3.8 Max there is no cheaper seller today. On Qwen3.8 27B the official seller is one of the more expensive ones. Neither fact is visible from Alibaba's pricing page, and both can change by the month. The Mercatus Token Index tracks where these models actually trade, and How a token exchange works covers what it takes to turn eighteen rate cards into one price.
Frequently asked questions
How much does the Qwen API cost?
As of September 2026, Qwen3.8 Max costs $2.00 per million input tokens and $6.00 per million output tokens on Alibaba Cloud Model Studio, with cached input at $0.25. Qwen3.8 Flash costs $0.15 and $0.47. The open-weight Qwen3.8 27B ranges from $0.15 to $0.45 input and $2.00 to $3.20 output across 13 third-party hosts. Older generations stay available: Qwen3.7 Max at $1.475 and $4.425, Qwen3.7 Flash at $0.03 and $0.13.
Is Qwen3.8 Max open-weight?
The Max endpoint is proprietary, but Alibaba released Qwen3.8 2.4T A95B on August 12, 2026 as the open-weight variant of the same architecture. Seven hosts serve it, six of them at the same $2.00 and $6.00 as the official Max endpoint.
What is the cheapest Qwen model?
Qwen3.7 Flash at $0.03 input and $0.13 output per million tokens. In the current generation, Qwen3.8 Flash at $0.15 and $0.47.
Who is the cheapest host for Qwen3.8 27B?
Darkbloom at $0.15 input and $2.00 output on September 7, 2026, followed by Reka AI, Parasail, and AkashML. Alibaba's official API at $0.425 is tenth of thirteen on input.
Is Qwen cheaper than DeepSeek?
Not at the flagship tier. Qwen3.8 Max costs $2.00 and $6.00 against DeepSeek V4 Pro's $0.66 and $1.98 off-peak, about 3x more. At the small-model tier they are close: Qwen3.8 Flash at $0.15 and $0.47 against DeepSeek V4 Flash at $0.22 and $0.66 official, or $0.0679 and $0.168 from the cheapest DeepSeek host. Qwen3.7 Flash at $0.03 and $0.13 undercuts everything.
Does Qwen have peak and off-peak pricing?
No. As of September 2026 all Qwen rates are flat around the clock. Among the models tracked on Mercatus, DeepSeek is the only one billing by time of day.
Why do all the hosts charge the same for Qwen3.8 2.4T?
The open weights launched nine days after the closed flagship at the flagship's exact price, and in the first month no host has undercut it. Whether that reflects serving cost for a 2.4T-parameter model or launch-window anchoring will show in whether the spread opens over the next month.
Methodology
Official Qwen rates are Alibaba Cloud Model Studio international-region list prices as reflected on OpenRouter's Alibaba Cloud International provider listings on September 7, 2026, including cache-read rates; China-region pricing may differ and was not reviewed. Third-party host prices, throughput, and uptime for Qwen3.8 2.4T A95B and Qwen3.8 27B are from OpenRouter's model pages on the same date and reflect list prices including any promotions in effect. Prices for older Qwen generations are OpenRouter base listings on the same date. Release dates are from OpenRouter listing dates and Alibaba's Qwen3.8 Max availability announcement of August 3, 2026. Workload calculations assume uncached input; the cached scenario assumes a 60 percent cache hit rate on input. DeepSeek, GLM-5.2, and Kimi K3 comparison figures are from the corresponding Mercatus pricing pages.
