Qwen3.8 Max Prime costs $3.301 per million input tokens and $9.902 per million output tokens on Alibaba Cloud's Beijing endpoint, against $1.65 and $4.951 for standard Qwen3.8 Max. That is a 2.0x multiple on both sides for the same 2.4 trillion parameter model, the same 1M context window and, in Alibaba's words, the same capabilities and restrictions. What the extra dollar buys is throughput: Alibaba says Prime raises tokens per second to 1.5 to 2 times the standard API.
Alibaba announced the tier at its Yunqi conference on September 22, 2026, and it appeared on OpenRouter's model list the next day with one provider. Standard Qwen3.8 Max is served by seven hosts. Prime is served by one.
Qwen3.8 Max Prime official pricing
Alibaba prices Model Studio by region, and the international (Singapore) rate card runs about 20 percent above Beijing for the same model.
| Region | Model | Input, $ per 1M | Output, $ per 1M | Prime multiple |
|---|---|---|---|---|
| Beijing | qwen3.8-max | $1.65 | $4.951 | |
| Beijing | qwen3.8-max-prime | $3.301 | $9.902 | 2.0x |
| Singapore (international) | qwen3.8-max | $2.00 | $6.00 | |
| Singapore (international) | qwen3.8-max-prime | $4.00 | $12.00 | 2.0x |
Beijing figures are from Alibaba Cloud's Model Studio pricing page. The Singapore Prime figure is Alibaba Cloud International's listing as shown on OpenRouter; standard Singapore rates are from Alibaba's own page. Context caching discounts and the 50 percent batch discount apply to standard Qwen3.8 Max. Alibaba's pricing page does not list either discount for the Prime model.
What Prime is
Prime is not a new model. Alibaba's documentation describes it as a fast mode for "output-speed-sensitive scenarios" where "TPS is raised to 1.5 to 2 times that of the standard API," and states that "model-supported capabilities and usage restrictions are the same as the original model." Same weights, same 1M context, same 131,072 token max output. The documentation also says that if the platform has spare resources, standard requests are not throttled, which is the honest description of what Prime sells: a place at the front of the queue when resources are not spare.
That makes Prime a priority tier, the inference equivalent of a reserved lane. The buyer pays double for a scheduling guarantee, and the guarantee is worth exactly as much as the queue is long. On a quiet day the standard tier may run at Prime speed for half the price. On a busy day the gap is the whole point.
The price of speed on a real workload
Take a production agent doing 10 million input tokens and 1 million output tokens a day, run on the Beijing endpoint.
| Standard Qwen3.8 Max | Qwen3.8 Max Prime | Difference | |
|---|---|---|---|
| Input, 10M tokens | $16.50 | $33.01 | $16.51 |
| Output, 1M tokens | $4.95 | $9.90 | $4.95 |
| Per day | $21.45 | $42.91 | $21.46 |
| Per year | $7,829 | $15,662 | $7,833 |
On the international rate card the same workload is $26.00 a day standard and $52.00 Prime, $9,490 a year apart. Caching narrows the standard bill further, since cache hits are discounted on the standard model and, per the pricing page, not on Prime.
So the question for a buyer is not whether Prime is fast. It is whether a 1.5 to 2x speedup on output is worth $7,800 to $9,500 a year on this workload, and whether the speed is needed at every hour or only at peak.
Prime against the open market
Qwen3.8 Max ships with open weights, and that changes the shape of this decision. Seven hosts serve the standard model on OpenRouter, six of them at Alibaba's $2 and $6, which is why the Qwen API pricing spread across hosts sits at 1x, the tightest of any major open model. Prime is not open. It is a serving tier on Alibaba's infrastructure, so no third-party host can offer it, and no third-party host can undercut it.
That leaves three ways to buy the same model:
- Standard Qwen3.8 Max from Alibaba or six other hosts at $2 and $6 international, with caching and batch discounts, and whatever throughput the host has spare.
- Qwen3.8 Max Prime from Alibaba only at 2x, with a throughput commitment.
- A third-party host with its own fast tier. Several hosts advertise speed on the standard weights. None of them carried a Prime-equivalent price on day one, which means the market has not yet priced throughput separately for this model. Alibaba has.
The spread on Qwen3.8 Max just went from 1x to 2x, and the whole 2x is one seller's throughput premium.
What it means for buyers
Prime is the clearest example yet of a provider unbundling capacity from tokens. The weights are free, the tokens cost $2, and the guarantee that your tokens arrive fast costs another $2. That premium is set by Alibaba and can move on Alibaba's schedule, the way DeepSeek's rates moved in August and the way GLM-5.3 Flash's launch price doubled when its promotion ended.
For a buyer with steady, speed-sensitive load, paying a 100 percent premium every month for capacity is the expensive way to lock it in. A forward contract on the Mercatus Forward Market reserves a month of inference capacity at a price set by the market rather than by one vendor, and can be resold if the load does not materialize. Prime cannot be resold. It can only be paid for again next month.
For everyone else, the standard tier at $2 and $6 with caching is still one of the cheapest frontier-class rates on the Token Index, and the 2x line above it is now the number to watch. If other hosts start charging for throughput on the same weights, the Qwen3.8 Max spread widens for the first time. If they do not, Alibaba's premium stands alone.
Frequently asked questions
What does Qwen3.8 Max Prime cost?
$3.301 per million input tokens and $9.902 per million output tokens on Alibaba Cloud's Beijing endpoint, and $4 and $12 on the international endpoint. Both are exactly 2x the standard Qwen3.8 Max rate in the same region.
Is Qwen3.8 Max Prime a different model?
No. Alibaba's documentation says capabilities and restrictions are the same as the original model. Prime is a serving tier with 1.5 to 2x the output tokens per second of the standard API.
Can I get Qwen3.8 Max Prime from a third-party host?
Not as of September 30, 2026. Prime is available only from Alibaba Cloud Model Studio. The standard model is served by seven hosts.
Is Prime worth it?
Only if output speed is the constraint and the standard tier is measurably slower on your traffic. On a 10M in / 1M out daily workload the premium is about $7,800 a year in Beijing and $9,500 internationally.
Does Prime get the caching and batch discounts?
Alibaba's pricing page lists context caching and batch discounts for standard Qwen3.8 Max and does not list them for Prime.
Methodology
Prices are from Alibaba Cloud Model Studio's model pricing page (Beijing region rows for qwen3.8-max and qwen3.8-max-prime, Singapore rows for qwen3.8-max) and from OpenRouter's listing of Alibaba Cloud International for qwen3.8-max-prime, both read September 30, 2026. Throughput and capability statements are Alibaba's own documentation language. Workload math uses cache-miss input rates, no batch discount, 365 days. Host counts are from OpenRouter on the same date. Prices in this category change on days' notice; verify before committing spend.
