DeepSeek V4 Pro costs $0.435 per million input tokens and $0.87 per million output tokens on the official DeepSeek API. Sixteen providers host it, and here is the unusual part: the official API is the cheapest of all of them. Third-party hosts charge anywhere from $0.59 to $1.74 per million input tokens for the same model, up to 4x the developer's own rate.
That is the exact opposite of the pattern on GLM-5.2, where the official API anchors the expensive end and discounters undercut it by 95 percent. Same market, same year, inverted structure. This page covers the full provider table, the cache discount that makes the official rate even harder to beat, the V4 Flash budget sibling, and how V4 stacks against V3.2.
What DeepSeek V4 is, briefly
Released April 24, 2026, DeepSeek V4 Pro is the flagship successor to the V3 line: an open-weight model with a 1M-token context window (up from V3.2's 163K), tool calling, and prompt caching, serving at roughly 63 tokens per second with a 1.33s time to first token. On public benchmark aggregates it sits in the 80th percentile for general intelligence.
It ships alongside DeepSeek V4 Flash, a non-reasoning variant at $0.09 input and $0.18 output with the same 1M window, one of the cheapest large-context models available from anyone.
DeepSeek V4 official API pricing
| Token type | V4 Pro | V4 Flash |
|---|---|---|
| Input $/1M | $0.435 | $0.09 |
| Input, cache hit $/1M | $0.004 | $0.0028 |
| Output $/1M | $0.87 | $0.18 |
The cache line deserves a second look: $0.004 against $0.435 is a 99 percent discount on cached input, the deepest cache discount we have documented on any major model. For agent and RAG workloads with heavy prompt reuse, this term alone decides the provider question, as the math below shows.
DeepSeek V4 Pro pricing by provider: all 16 hosts
| Provider | Input $/1M | Output $/1M | Cached input $/1M |
|---|---|---|---|
| DeepSeek (official) | $0.435 | $0.87 | $0.004 |
| GMI Cloud | $0.592 | $1.183 | $0.049 |
| StreamLake | $0.609 | $1.218 | $0.051 |
| Novita | $0.632 | $1.263 | $0.053 |
| OpenRouter (routed) | $0.632 | $1.263 | $0.053 |
| Cloudflare | $1.15 | $2.55 | $0.20 |
| DeepInfra | $1.30 | $2.60 | — |
| Alibaba | $1.416 | $2.832 | $0.118 |
| SiliconFlow | $1.502 | $3.135 | $0.135 |
| Venice | $1.65 | $3.301 | $0.33 |
| AtlasCloud | $1.68 | $3.38 | $0.13 |
| Parasail | $1.74 | $3.48 | $0.10 |
| Together | $1.74 | $3.48 | — |
| BaseTen | $1.74 | $3.48 | $0.145 |
| Fireworks | $1.74 | $3.48 | $0.145 |
| WandB | $1.74 | $3.48 | $0.14 |
The structure is a clean inversion of the GLM-5.2 table. There, the official API held the high end while discounters raced to $0.07. Here, nobody beats the developer, the nearest third party is 36 percent more expensive, and a cluster of six hosts sits at exactly 4x official list.
Why the developer is the cheapest host this time
Two markets, opposite structures, one explanation: who is trying to win the traffic.
Z.AI prices GLM-5.2 at a premium and lets third parties compete for volume. DeepSeek does the reverse: it prices its own API at or near serving cost and treats third-party hosts as overflow. For a lab that runs its own highly optimized inference stack, underpricing every reseller is a deliberate strategy, it keeps the traffic, the usage data, and the customer relationship at home. Third parties serving V4 Pro carry their own GPU economics plus margin, and without a price advantage they compete on other grounds: enterprise terms, data residency, platform integration, bundling.
The lesson for buyers is that "official vs third-party" has no fixed answer. It flips model by model, which is why rate-card assumptions ported from one model to another keep being wrong, the pattern documented across the whole market in Why Token Prices Differ.
What a real workload pays
A production assistant doing 10M input and 1M output tokens per day on V4 Pro:
| Provider | Daily | Monthly | Annual |
|---|---|---|---|
| DeepSeek official | $5.22 | $156.60 | $1,905 |
| Mid-tier host (Cloudflare, $1.15/$2.55) | $14.05 | $421.50 | $5,128 |
| Top-tier host ($1.74/$3.48) | $20.88 | $626.40 | $7,621 |
Now add a 60 percent cache hit rate on input:
| Provider | Daily with 60% cached input |
|---|---|
| DeepSeek official ($0.004 cached) | $2.63 |
| Best third-party cache (GMI, $0.049) | $3.85 |
The official API's lead widens from 12 percent (against the nearest competitor uncached) to 25 percent or more with caching, and against the $1.74 tier the gap approaches 8x. There is no cache mix at which a third party wins on price for this model. That is rare, and it is worth saying plainly: for pure cost, V4 Pro's answer is the official API, full stop. The remaining reasons to pay a third party are compliance, latency, region, or platform, all legitimate, none of them price.
And if the workload doesn't need reasoning: V4 Flash at $0.09/$0.18 runs the same 10M/1M day for $1.08, with the same million-token window.
DeepSeek V4 vs V3.2 vs GLM-5.2
| V4 Pro | V3.2 | GLM-5.2 | |
|---|---|---|---|
| Official input / output $/1M | $0.435 / $0.87 | $0.28 / $0.42 | $1.40 / $4.40 |
| Best available input $/1M | $0.435 (official) | ~$0.20 (3rd party) | $0.07 (3rd party) |
| Official cached input $/1M | $0.004 | $0.028 | $0.26 |
| Context window | 1M | 163K | 1M |
| Cheapest host | The developer | Third party | Third party |
| Cross-host spread | 4x | ~3x | 30x |
Three models, three completely different market structures. The only reliable way to buy is to run your own token mix against current rates, which is what the blended benchmarks on the Mercatus Token Index exist for. For the V3.2 breakdown, see DeepSeek V3.2 API Pricing.
How to choose a V4 provider
- Cost, any workload shape: official API. Cheapest uncached, untouchable cached. This model is the exception where the answer is simple.
- Long-context workloads: the 1M window is native; confirm third parties serve the full window before paying their premium for other reasons.
- No reasoning required: V4 Flash at $0.09/$0.18 is the strongest budget option in the 1M-context class from any developer.
- Compliance, region, or platform needs: the 4x premium at the top of the table is what those requirements currently cost on this model. Price it explicitly rather than absorbing it silently.
- Recheck monthly. V3.2's market compressed and GLM's collapsed within a quarter of launch; V4's structure is four months old and will move too.
Frequently asked questions
How much does DeepSeek V4 cost per million tokens?
V4 Pro: $0.435 input and $0.87 output on the official API, with cached input at $0.004. V4 Flash: $0.09 input and $0.18 output. Third-party hosts charge up to $1.74 input for V4 Pro.
Who is the cheapest DeepSeek V4 provider?
The official DeepSeek API, at $0.435 per million input tokens. It is the lowest of all 16 hosts, and its 99 percent cache discount widens the lead on cache-heavy workloads.
Why do third parties charge more than the official API?
DeepSeek prices its own API near serving cost to keep traffic first-party. Third-party hosts carry their own GPU costs plus margin, so without a price edge they compete on enterprise terms, regions, and platform features instead.
What is the difference between V4 Pro and V4 Flash?
Pro is the flagship with reasoning capability at $0.435/$0.87. Flash is the non-reasoning variant at $0.09/$0.18. Both carry the 1M-token context window.
Is DeepSeek V4 cheaper than V3.2?
Not at list: V3.2 is $0.28/$0.42 official against V4 Pro's $0.435/$0.87. V4 buys the 1M context window and a far deeper cache discount ($0.004 vs $0.028), so cache-heavy and long-context workloads often run cheaper on V4 despite the higher list price.
How does V4's provider spread compare to other models?
About 4x from cheapest to most expensive host, versus roughly 3x on V3.2 and 30x on GLM-5.2. Every model's hosting market has its own structure, which is why per-model pricing pages exist.
Methodology
Provider pricing compiled from published rate cards and public pricing trackers (sourced from OpenRouter and Helicone data), August 2026. Benchmark percentiles from public aggregates. Workload examples use stated token volumes and cache assumptions; recompute with your own mix. V3.2 and GLM-5.2 comparison figures from the corresponding Mercatus pricing pages. Last verified: 2026-08-12.
