DeepSeek V4 Pro now costs $0.66 per million input tokens off-peak and $1.32 at peak, with output at $1.98 and $3.96, following the price increase that took effect August 16, 2026 at 16:00 UTC. The flat-rate era ended with it: DeepSeek now bills by time of day, with off-peak set at half the peak rate. Until the change, V4 Pro ran at $0.435 input and $0.87 output, rates that survived the model's exit from preview on August 12 by only four days.
Those rates have a hard end date. On September 10, 2026, DeepSeek launched V4.1 Flash and announced that from 04:00 UTC on September 14, every request to deepseek-v4-pro will be routed to V4.1 Flash and billed at V4.1 Flash rates, $0.15 input and $0.60 output off-peak, until a V4.1 Pro is released. V4 Pro is leaving the official API less than a month after it was repriced. The new model and the routing are covered in DeepSeek V4.1 Flash API pricing. This page covers the V4 Pro rates, the August increase, and what a repricing this sharp says about how token prices actually work. The short version: the cheapest flagship API on the market raised prices up to 1,100 percent on ten days' notice, then retired the model four weeks later, and buyers had no instrument to lock either the rate or the product.
DeepSeek V4 Pro official API pricing
| V4 Pro off-peak | V4 Pro peak | V4.1 Flash off-peak | V4.1 Flash peak | |
|---|---|---|---|---|
| Input, cache miss $/1M | $0.66 | $1.32 | $0.15 | $0.30 |
| Input, cache hit $/1M | $0.022 | $0.044 | $0.003 | $0.006 |
| Output $/1M | $1.98 | $3.96 | $0.60 | $1.20 |
| Context window | 1M tokens | 1M tokens | 1M tokens | 1M tokens |
| Max output | 384K tokens | 384K tokens | 384K tokens | 384K tokens |
Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday through Friday; everything else, including weekends, bills off-peak at half the peak rate. The cache-hit discount survived in relative terms, cache hits still cost about 3 percent of a cache miss, but the absolute cached rate rose more than 10x, from $0.003625 to $0.022 off-peak. The 1M context window still carries no surcharge.
Both models support JSON output, tool calls, and Anthropic API compatibility. V4.1 Flash replaced V4 Flash on the official API on September 10, 2026. The V4 Pro columns are live until September 14.
What changed at GA
On price, nothing. The 0813 build is the first flagship DeepSeek release to exit preview with its preview pricing intact.
On the model itself, V4 Pro is a mixture-of-experts system with 1.6 trillion total parameters and 49 billion active per token. DeepSeek reports that its compressed attention design cuts inference compute to 27 percent of V3.2's at maximum context, with KV cache requirements at 10 percent. Those two engineering numbers are the economic story: they are why a 1M-context flagship can list at $0.435 at all. DeepSeek also reports benchmark results for the Pro Max mode including 80.6 percent on SWE-bench Verified and 93.5 percent on LiveCodeBench.
For how infrastructure efficiency translates into token prices generally, see why token prices differ across providers.
The DeepSeek V4 Pro price increase: what changed on August 16
The new rates took effect August 16, 2026 at 16:00 UTC, replacing the single flat rate with peak and off-peak billing. V4 Pro input on a cache miss went from $0.435 to $0.66 off-peak and $1.32 at peak. Output went from $0.87 to $1.98 and $3.96. Across the V4 line, the increases range from roughly 50 percent to over 1,100 percent depending on model, token type, and time of day.
This is worth sitting with, because it is a clean illustration of how posted prices work. A posted price is not a market price. It is a number one seller chooses, and it can move by announcement, in either direction, at any time. DeepSeek cut V3.2 prices repeatedly through 2025 and 2026. It has now made the opposite move, from a position where its official API is already among the cheapest hosts of its own model.
For buyers, the practical readings are now concrete. Cost models built on $0.435 input broke overnight; the same workload now costs 64 percent more off-peak and over 3x at peak. Cache-heavy workloads were hit hardest, since the near-free cached rate rose more than 10x. And third-party hosts, who set prices independently, re-sorted within three weeks: the cheap ones (GMI Cloud, StreamLake, Novita) repriced upward toward the new official peak, while DigitalOcean entered at $0.87 input and $1.74 output, below the official peak on both sides and below off-peak on output. The full provider table is on the DeepSeek V4 API pricing page.
What a real workload pays now
Take a production workload of 10M input tokens and 1M output tokens per day.
// text
Same workload, 10M input + 1M output tokens per day:
Before August 16 (flat rate):
No caching: $5.22/day ($1,905/year)
60% cache hits: $2.63/day ($961/year)
After August 16 (off-peak):
No caching: $8.58/day ($3,132/year)
60% cache hits: $4.75/day ($1,734/year)
After August 16 (peak):
No caching: $17.16/day ($6,263/year)
60% cache hits: $9.50/day ($3,468/year)The same workload now costs 64 percent more if it runs entirely off-peak, and 3.3x more at peak. The spread between the best case (off-peak, cached) and worst case (peak, uncached) is now wider than the entire old bill. Two lessons survived the reprice: the blended rate is still what you actually pay, and as the V3.2 analysis showed, cache terms still dominate the blend. What's new is that the clock now matters too: scheduling batch work off-peak is worth an automatic 50 percent discount.
V4 Pro across providers
Third-party hosts list V4 Pro between $0.87 and $1.91 per million input tokens as of September 2, 2026. Off-peak, the official API at $0.66 is still the cheapest host on input and cached tokens; only DigitalOcean ($0.87 / $1.74) beats it on output. At peak ($1.32 / $3.96), seven hosts undercut the official API on input and all sixteen on output. The full 16-provider table is on the DeepSeek V4 API pricing page, and every DeepSeek model and host is on DeepSeek API pricing. After September 14 those third-party hosts are the only place V4 Pro is sold. The official API stops serving it, so a buyer who needs the V4 Pro weights specifically, rather than whatever DeepSeek routes deepseek-v4-pro to, has to buy from a host.
The re-sort happened fast. GMI Cloud, StreamLake, and Novita sat at $0.59 to $0.63 before the increase and moved to $1.03 to $1.60 within three weeks, tracking the new official peak rather than the old flat rate. On V4 Flash the opposite happened: most hosts held DeepSeek's old $0.14 / $0.28 price, which left the official API as the second most expensive Flash host in the market. Which prices move, and when, is exactly the kind of question a posted-price system answers slowly and a live market answers continuously.
The window, in hindsight
The decision window lasted ten days from DeepSeek's first warning on August 6, and only three days from August 13, when the actual rates were published, to August 16, when they took effect. During it, there was no way to lock the old economics: the official API sells at posted prices with no term structure, so buyers could only watch the date approach. Anyone budgeting a year of V4 Pro at August 15 rates woke up to a bill that is 64 to 230 percent higher. The routing notice repeated the pattern. DeepSeek said on September 10 that V4 Pro traffic moves to V4.1 Flash on September 14. Four days, and a different model at the end of them.
This is the structural gap forward markets exist to fill. When a seller can reprice by announcement, a buyer with a forward curve can price that risk and lock delivery months ahead; a buyer facing a posted price cannot. Physically settled token forwards are the instrument built for exactly this, and repricings like this one are the reason demand for them is growing. The Token Index tracks where V4 Pro and other open-weight models actually trade as the market re-sorts.
Frequently asked questions
What does DeepSeek V4 Pro cost right now?
As of August 16, 2026: $0.66 per million input tokens off-peak and $1.32 at peak (cache miss), $0.022 to $0.044 on cache hits, and $1.98 to $3.96 per million output tokens. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday. The 1M context window carries no surcharge. From September 14, 2026, deepseek-v4-pro requests on the official API are served by V4.1 Flash and billed at $0.15 input and $0.60 output off-peak.
When did prices increase, and by how much?
August 16, 2026 at 16:00 UTC. V4 Pro input went from $0.435 to $0.66 off-peak and $1.32 peak; output from $0.87 to $1.98 and $3.96. Cached input rose more than 10x. DeepSeek also moved from a flat rate to peak/off-peak billing, with off-peak at half of peak. DeepSeek warned of an increase on August 6 without figures and published the rates on August 13, three days before they took effect.
Who is the cheapest DeepSeek V4 Pro host after the increase?
Off-peak, the official DeepSeek API at $0.66 input and $0.022 cached. At peak, DigitalOcean at $0.87 input and $1.74 output, then StreamLake and Baidu Qianfan at about $1.03 / $2.06. For a workload that runs around the clock it is roughly a tie between the official API and DigitalOcean at about $10.40 a day on 10M input and 1M output.
Is V4 Pro GA pricing different from preview pricing?
No. The 0813 GA build kept preview pricing unchanged.
Should I use V4 Pro or V4 Flash?
V4 Flash was retired on September 10, 2026. Its replacement, V4.1 Flash, lists at $0.15/$0.60 off-peak with the same 1M context window, and DeepSeek says it beats V4 Pro on performance, speed, and cost, which is why V4 Pro traffic routes to it from September 14. Test V4.1 Flash on your actual task before then. If it passes, the routing does the rest. If it fails, third-party hosts still serve the V4 Pro weights.
Methodology
Official pricing verified against DeepSeek's API documentation on August 13, 2026. Release details from DeepSeek's V4 Pro 0813 general availability announcement. Post-increase rates verified against DeepSeek's API documentation. Third-party provider range from OpenRouter's V4 Pro page on September 2, 2026. V4.1 Flash rates and the V4 Pro routing rule verified against DeepSeek's API documentation and release notice on September 10, 2026. Last verified: 2026-09-10.
