HomeBlogDeepSeek V4.1 Flash API Pricing: Rates, Hosts, V4 Pro Routing
GeneralSep 10, 202611 min read

DeepSeek V4.1 Flash API Pricing: Rates, Hosts, V4 Pro Routing

DeepSeek V4.1 Flash launched September 10, 2026 at $0.15 input and $0.60 output off-peak. V4 Flash is retired and every V4 Pro request routes to V4.1 Flash from September 14. Official rates, day-one host prices, and what the routing means for buyers.

M

Mercatus Compute

Author

DeepSeek V4.1 Flash API Pricing: Rates, Hosts, V4 Pro Routing

DeepSeek V4.1 Flash went live on the official API on September 10, 2026 at $0.15 per million input tokens and $0.60 per million output tokens during off-peak hours, double that at peak, with cache hits at $0.003 off-peak. The model name is deepseek-flash. It replaces V4 Flash outright, and from September 14 it also replaces V4 Pro: every request to deepseek-v4-pro will be served by V4.1 Flash and billed at V4.1 Flash rates until a V4.1 Pro exists.

That second part is the news for anyone running a DeepSeek workload in production. In a single notice, DeepSeek retired two models and rerouted a third onto a new architecture, with four days between the announcement and the switch. This page covers the official rate card, how it compares with the models it replaces, what the four third-party hosts serving V4.1 Flash on launch day are charging, and the workload math for buyers deciding whether to stay on the official API or move.

DeepSeek V4.1 Flash official API pricing

Two models are listed on the official API as of September 10, 2026. Both have a 1M-token context window and a 384K maximum output. Prices are USD per million tokens and took effect at 04:00 UTC on September 10.

ModelInput, cache miss (off-peak / peak)Input, cache hit (off-peak / peak)Output (off-peak / peak)VisionConcurrency
deepseek-flash (V4.1 Flash)$0.15 / $0.30$0.003 / $0.006$0.60 / $1.20Yes2,500
deepseek-v4-pro (V4-Pro-0813)$0.66 / $1.32$0.022 / $0.044$1.98 / $3.96No500

Peak hours are unchanged: 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday through Friday, with everything else including weekends billed at the off-peak rate. The time-zone breakdown is in DeepSeek peak and off-peak hours.

Three notes on the rate card. The legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted, but the models behind them are retired; requests to either name are served by V4.1 Flash and billed at the Flash price. The V4 Pro row is live for four more days. And the concurrency limit on Flash is five times Pro's, which matters for agent fleets more than the per-token price does.

What V4.1 Flash is. DeepSeek describes it as the smallest model in a new architecture family: 552 billion parameters in a mixture-of-experts layout, with a causal encoder-decoder design that activates 8 billion parameters on input and 16 billion on output. Native image understanding is built in rather than bolted on as it was in the V4-Flash-Vision-Exp model. DeepSeek says the KV cache needs a quarter of the HBM and an eighth of the SSD storage of the previous generation, and that its own and third-party tests put V4.1 Flash ahead of V4 Pro on performance, cost, speed, and total task time. The weights are published on Hugging Face with a technical report.

How the new rate card compares with V4 Flash and V4 Pro

V4.1 Flash is a new model on a new cost curve, not a repricing of an existing one. The table puts its rates next to the two models it replaces, using the off-peak column for all three.

Token typeV4 Flash (retired)V4 Pro (routed from Sept 14)V4.1 Flash
Input, cache miss$0.22$0.66$0.15
Input, cache hit$0.007$0.022$0.003
Output$0.66$1.98$0.60

The cache-hit rate is the one to look at. DeepSeek's stated reason for the lower cache price is the smaller KV cache, and cache hits are where agent workloads spend most of their input budget. On the official endpoint, OpenRouter's monitoring shows a 93.9 percent cache hit rate for V4.1 Flash in its first hours, which is the kind of traffic that makes the $0.003 rate the number that decides the bill.

The comparison that matters commercially is with V4 Pro, because that is the traffic being moved. A V4 Pro buyer paying $0.66 and $1.98 off-peak will, from September 14, be billed $0.15 and $0.60 for the same requests, served by a different model that DeepSeek says is better. Whether it is better for a specific workload is something each buyer has to test in four days.

The V4 Pro routing and what it means for buyers

The mechanics, from DeepSeek's pricing page: from 12:00 Beijing time (04:00 UTC) on September 14, 2026, all requests to deepseek-v4-pro are routed to V4.1 Flash and billed at V4.1 Flash rates. This continues until V4.1 Pro is released. DeepSeek is phasing V4 Pro out because, in its words, V4.1 Flash "has comprehensively surpassed V4 Pro in performance, cost, speed, and total time."

For most Pro buyers this is a lower bill with no code change. For buyers who pinned V4 Pro for a reason, whether an eval they passed on Pro, a compliance sign-off on a specific model version, or a reasoning task they have not yet validated on the new architecture, the official API offers no way to keep Pro after September 14. The routing is not opt-in.

The third-party market is the continuity option. V4-Pro-0813 weights are open, and on September 10 OpenRouter listed hosts serving them at $0.5808 input and $1.742 output and at $0.66 and $1.98, the same rates as the official API off-peak. A buyer who needs V4 Pro specifically can keep it, just not from DeepSeek. That is the same structure the August 16 increase produced for V4 Flash, where sixteen hosts kept the pre-increase price after DeepSeek moved, covered in DeepSeek V4 Pro API pricing.

The larger point is about who controls the product. One seller decided, on September 10, which model its Pro customers would be running on September 14. There was no market to price the transition against and no contract that held the old terms. Many sellers serving the same open weights is what turns that from a decision buyers absorb into one they can route around.

DeepSeek V4.1 Flash pricing by host, launch day

Four hosts served V4.1 Flash on OpenRouter on September 10, 2026, the day of release. Prices are per million tokens. Throughput and uptime are OpenRouter's measurements over the same day.

HostInputOutputCache readThroughputUptime
DeepSeek official (off-peak)$0.15$0.60$0.003139 tps99.99%
DeepSeek official (peak)$0.30$1.20$0.006139 tps99.99%
NovitaAI$0.30$1.20$0.006121 tps100%
DeepInfra$0.30$1.20$0.0065 tps96.69%
Venice$0.375$1.50$0.00756 tps77.54%

On launch day the official API is the cheapest place to buy V4.1 Flash at any hour. Novita and DeepInfra list at exactly DeepSeek's peak rate, which is 2x the official off-peak price and matches it only during the 35 peak hours a week. Venice is 2.5x. Two of the three third parties were still ramping, at 5 and 6 tokens per second against 139 on the official endpoint.

This is the opposite of where the V4 Flash market ended up. On September 2, sixteen of seventeen hosts listed V4 Flash below DeepSeek's own off-peak rate, and the cheapest was at a third of it. The difference is time. V4 Flash weights had been out for months and hosts had optimized their serving stacks and margins around them. V4.1 Flash is hours old, with a new architecture that inference engines have not yet tuned for, and hosts have anchored on DeepSeek's peak number while they work out their own costs.

The next few weeks will show whether the pattern repeats. If the third-party market on V4.1 Flash goes the way of V4 Flash, the official rate stops being the floor. If the new architecture is harder to serve efficiently than DeepSeek's own stack allows, the official API stays cheapest for longer. Either way, the launch-day table is a snapshot, not a settled market, which is why the Mercatus Token Index tracks the spread rather than any one host.

What a real workload pays

A production assistant doing 10 million input tokens and 1 million output tokens a day, uncached. Routes are priced at list rates on September 10, 2026. V4 Flash and V4 Pro rows use official off-peak rates for the models being replaced, plus the cheapest third-party host still serving the old V4 Flash weights.

RouteDailyAnnual
V4.1 Flash, official, all off-peak$2.10$767
V4.1 Flash, official, uniform traffic (21% peak)$2.54$927
V4.1 Flash, official, all peak$4.20$1,533
V4.1 Flash, Novita or DeepInfra$4.20$1,533
V4.1 Flash, Venice$5.25$1,916
V4 Flash, official off-peak (retired Sept 10)$2.86$1,044
V4 Flash 0731 weights, cheapest host$0.66$241
V4 Pro, official off-peak (routed Sept 14)$8.58$3,132

Three things fall out of the table.

For V4 Pro traffic, the routing is the whole story. The same 10M/1M workload that costs $8.58 a day on V4 Pro off-peak will cost $2.10 on V4.1 Flash off-peak from September 14, on the official API, with no change to the code. A buyer who validates the new model on their task in the next four days gets that number. A buyer who cannot, and moves Pro traffic to a third-party host to keep the old model, pays $8.58 or more.

For V4.1 Flash, buy it from DeepSeek for now. A round-the-clock workload pays $2.54 a day on the official API and $4.20 on the cheapest third party. Even a workload that runs entirely in peak hours only reaches parity with Novita and DeepInfra. There is no third-party discount on this model yet.

Cache changes the answer more than the clock does. At a 60 percent cache hit rate, the official off-peak route falls to $1.22 a day, or $445 a year, because 6 million of the 10 million input tokens bill at $0.003 instead of $0.15. The peak and off-peak gap on cached tokens is $0.003, small enough that a cache-heavy agent barely notices the schedule. An uncached workload is where the 2x peak multiplier still bites.

One more row deserves a look. The old V4 Flash 0731 weights are still served by third-party hosts, and the cheapest listed at $0.05 input and $0.16 output on September 10, which runs the workload for $0.66 a day. DeepSeek retired the model; the market did not. For a buyer whose task was fine on V4 Flash and who has no need for the new architecture, the retired model from a third party is a quarter of the price of the new one from DeepSeek. That is what a competitive supply side does when the developer moves on.

Which DeepSeek model and host to use

Running V4 Pro today: test V4.1 Flash on your workload before September 14. If it passes, do nothing; the routing moves you and the bill follows. If it fails, or you need a pinned model version, move to a third-party V4 Pro host now rather than discovering the switch in production.

Running V4 Flash today: you are already on V4.1 Flash. The legacy name routes there and bills at the new rate. Check output quality and the vision behavior if your pipeline sends images, since the encoder is different.

Starting fresh on V4.1 Flash: use the official API. It is the cheapest host at every hour on launch day and the fastest by a wide margin. Revisit the host table in two to three weeks.

Cost-sensitive and happy with the old model: V4 Flash 0731 weights from a third-party host, from $0.05 and $0.16. That market will thin out as hosts move capacity to V4.1, so it is not a long-term plan.

Cache-heavy agents: the official $0.003 cache-hit rate is the lowest in the DeepSeek line and half of any third party's on this model. Structure prompts so the shared prefix hits.

The model-by-model comparison against GLM, Qwen, and Kimi is in the LLM API pricing comparison.

Where DeepSeek pricing goes from here

DeepSeek's last two pricing events point in opposite directions and make the same case. In August, the official rate went up and the third-party market held the old price, so the official API went from cheapest to among the most expensive within a week. In September, a new model arrived at a lower rate card and the third-party market, on day one, sits at double the official price. In both cases the posted price and the market price were different numbers, and only one of them was set by more than one seller.

What has not changed is how the decisions get made. A seller published a rate, retired two products, and rerouted a third, with a few days' notice and no instrument for buyers to hold their terms. Better models create competition, and the third-party market on the old weights is proof that buyers now have somewhere else to go. What they still lack is a single price that reflects all of it. How a token exchange works covers how that number gets set by trades instead of by a pricing page, and How to lock in token prices covers what a buyer can do about a routing notice like this one before it lands. The full DeepSeek line, official rates and every host, is kept current in DeepSeek API pricing.

Frequently asked questions

How much does DeepSeek V4.1 Flash cost?
As of September 10, 2026, the official DeepSeek API charges $0.15 per million input tokens and $0.60 per million output tokens off-peak, and $0.30 and $1.20 at peak. Cache hits are $0.003 off-peak and $0.006 at peak. The model name is deepseek-flash. Third-party hosts listed it at $0.30 to $0.375 input and $1.20 to $1.50 output on launch day.

What happened to DeepSeek V4 Flash?
It was retired on September 10, 2026. The names deepseek-v4-flash and deepseek-v4-flash-vision-exp still work but route to V4.1 Flash and bill at the V4.1 Flash rate. The old V4 Flash weights are still served by third-party hosts.

Is DeepSeek V4 Pro being discontinued?
Yes. From 04:00 UTC on September 14, 2026, all requests to deepseek-v4-pro on the official API are routed to V4.1 Flash and billed at V4.1 Flash rates, until V4.1 Pro is released. V4 Pro weights remain available from third-party hosts.

Is V4.1 Flash better than V4 Pro?
DeepSeek says its own and third-party tests put V4.1 Flash ahead of V4 Pro on performance, cost, speed, and total task time, and that is its stated reason for the routing. Buyers with Pro in production have until September 14 to confirm that on their own workloads.

Does V4.1 Flash support images?
Yes. Native image understanding is built into the model. The separate V4-Flash-Vision-Exp model is retired.

Are the V4.1 Flash weights open?
Yes. DeepSeek published the weights and a technical report on Hugging Face at release. Four hosts were serving them on OpenRouter within hours.

Is the official API the cheapest way to run V4.1 Flash?
On launch day, yes, at every hour. The three third-party hosts listed at 2x to 2.5x the official off-peak rate. That was also true of V4 Flash at launch, and by September 2026 sixteen hosts undercut the official rate, so the table is expected to move.

Methodology

Official V4.1 Flash and V4 Pro rates, peak-hour definitions, model specifications, the V4 Flash retirement, and the V4 Pro routing rule are from the DeepSeek API documentation pricing page and the September 10, 2026 release notice, read on the day of release. Retired V4 Flash rates are the official off-peak rates in force from August 16 to September 10, 2026. Third-party host prices, throughput, uptime, and cache hit rates for V4.1 Flash, and third-party prices for V4 Pro 0813 and V4 Flash 0731 weights, are from OpenRouter's model pages on September 10, 2026 and reflect list prices on that date. Workload calculations assume uncached input except where a cache rate is stated; the uniform-traffic case weights peak hours at 35 of 168 weekly hours. Launch-day host prices are expected to change as more providers list the model.