HomeBlogDeepSeek API Pricing: Official Rates vs Third-Party Hosts
GeneralSep 2, 20268 min read

DeepSeek API Pricing: Official Rates vs Third-Party Hosts

DeepSeek V4 Flash costs $0.22 to $1.32 per million tokens on the official API and as little as $0.07 through third-party hosts. Every DeepSeek model, every host, and what the August 16 price increase changed.

M

Mercatus Compute

Author

DeepSeek API Pricing: Official Rates vs Third-Party Hosts

DeepSeek's official API charges $0.22 per million input tokens and $0.66 per million output tokens for V4 Flash during off-peak hours, and exactly double that at peak. V4 Pro runs $0.66 input and $1.98 output off-peak, $1.32 and $3.96 at peak. Those are the rates in force since August 16, 2026, when DeepSeek raised output prices by up to 371 percent with the specific numbers published three days before they took effect.

The official page is no longer the whole story. Seventeen third-party hosts serve the same V4 weights, and since the increase most of them are cheaper than DeepSeek itself. The lowest listed V4 Flash rate is $0.0679 input and $0.168 output, which is one eighth of the official peak output price for an identical model. This page covers the official rates, the increase, every host's price, and the workload math that decides where a DeepSeek bill should actually be paid.

DeepSeek official API pricing

Three models are live on the official API as of September 2, 2026. All three have a 1M-token context window and a 384K max output. Prices are USD per million tokens.

ModelInput, cache miss (off-peak / peak)Input, cache hit (off-peak / peak)Output (off-peak / peak)
deepseek-v4-flash$0.22 / $0.44$0.007 / $0.014$0.66 / $1.32
deepseek-v4-pro$0.66 / $1.32$0.022 / $0.044$1.98 / $3.96
deepseek-v4-flash-vision-exp$0.22 / $0.44$0.007 / $0.014$0.66 / $1.32

Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday through Friday. Everything else, including weekends, bills at the off-peak rate. That is 35 peak hours out of 168 in a week, roughly 21 percent of the clock. In US terms, peak runs 9 p.m. to midnight and 2 a.m. to 6 a.m. Eastern, which is why the schedule hits Asian and European business hours far harder than American ones. The full time-zone breakdown is in DeepSeek peak and off-peak hours.

The vision model bills images as input tokens, up to 384 tokens per image, at V4 Flash rates. Cache hits apply when a request repeats a prefix the API has already processed, which is common in agent loops and long system prompts.

What changed on August 16

Before August 16, DeepSeek charged one flat rate around the clock. V4 Pro listed at $0.435 input and $0.87 output, a price it had held since May 23. V4 Flash listed at $0.14 input and $0.28 output.

Token typeOld flat rateNew off-peakNew peakChange at peak
V4 Flash input, cache miss$0.14$0.22$0.44+214%
V4 Flash output$0.28$0.66$1.32+371%
V4 Flash input, cache hit$0.0028$0.007$0.014+400%
V4 Pro input, cache miss$0.435$0.66$1.32+203%
V4 Pro output$0.87$1.98$3.96+355%
V4 Pro input, cache hit$0.003625$0.022$0.044+1,114%

Even the off-peak rates, the cheapest DeepSeek now offers, are 52 to 136 percent above the old flat price on regular tokens and 2.5 to 6 times higher on cached tokens.

The timeline matters as much as the numbers. DeepSeek posted a notice on August 6 saying a significant increase was coming, without figures. The rates were published on August 13 alongside the V4 Pro general-availability release. They took effect on August 16 at 16:00 UTC. A buyer who had budgeted on the August 12 price list had three days to react once the numbers were public, and no instrument to hedge with. DeepSeek's stated reason was to "allocate resources more reasonably" and push developer workloads toward less congested hours. The mechanics of that repricing, and what it cost a typical buyer, are worked through in How to lock in token prices.

DeepSeek V4 Flash pricing by host

Prices below are per million tokens as listed on OpenRouter on September 2, 2026, sorted cheapest first. Two hosts were running promotional discounts on that date, marked with an asterisk. The official DeepSeek API is included at both rates for comparison.

HostInputOutputCache read
DigitalOcean$0.0679$0.168$0.0168
StreamLake*$0.0886$0.177$0.0177
DeepInfra$0.09$0.18$0.018
GMI Cloud*$0.112$0.224$0.0224
SiliconFlow$0.13$0.28$0.028
Alibaba Cloud$0.134$0.268$0.0268
Venice$0.138$0.275$0.028
Novita$0.14$0.28$0.028
NextBit$0.14$0.28$0.028
AtlasCloud$0.14$0.28$0.028
Baidu Qianfan$0.14$0.28$0.028
CoreWeave$0.14$0.28$0.07
Parasail$0.14$0.28$0.07
Mancer$0.185$0.50n/a
Phala$0.20$0.40$0.07
Azure (US)$0.21$0.56$0.031
DeepSeek official, off-peak$0.22$0.66$0.007
Cloudflare$0.44$1.32$0.014
DeepSeek official, peak$0.44$1.32$0.014

The pattern is hard to miss. A block of six hosts sits at exactly $0.14 and $0.28, which is DeepSeek's own pre-August price list. The third-party market anchored on the old official rate and stayed there when DeepSeek moved. Only Cloudflare matched the new peak rate. The result is that the official API is now the second most expensive place to buy V4 Flash output at any hour of the day.

The spread from cheapest to most expensive host is 6.5x on input and 7.9x on output, for the same weights.

DeepSeek V4 Pro pricing by host

Same source and date. The Pro spread is narrower because Pro is more expensive to serve, but the official API's position has still changed.

HostInputOutputCache read
DeepSeek official, off-peak$0.66$1.98$0.022
DigitalOcean$0.87$1.74$0.174
StreamLake$1.027$2.055$0.086
Baidu Qianfan$1.029$2.058$0.085
GMI Cloud$1.044$2.088$0.087
Ionstream$1.131$2.262$0.094
CoreWeave$1.15$2.55$0.20
DeepInfra$1.30$2.60$0.10
DeepSeek official, peak$1.32$3.96$0.044
Alibaba Cloud$1.416$2.832$0.118
SiliconFlow$1.502$3.135$0.135
Novita$1.60$3.20$0.135
Venice$1.65$3.301$0.33
AtlasCloud$1.68$3.38$0.13
NextBit$1.72$3.45$0.14
Baseten$1.74$3.48$0.145
Parasail$1.74$3.48$0.10
Azure (US)$1.91$3.83$0.16

Off-peak, DeepSeek is still the cheapest V4 Pro host on input and cache reads, and only DigitalOcean beats it on output. At peak, DeepSeek drops to the middle of the table: seven hosts undercut it on input and all sixteen undercut it on output. Third-party output prices run $1.74 to $3.83, and the official peak rate of $3.96 is the highest number in the column.

The official API is no longer the default cheapest host

The tables above are list prices. What a buyer actually pays depends on the token mix and the clock. The workload below is 10 million input tokens and 1 million output tokens per day, uncached, which is a typical shape for a retrieval-heavy production app.

RouteV4 Flash, per dayV4 Flash, per yearV4 Pro, per dayV4 Pro, per year
Official, old flat rate (pre-Aug 16)$1.68$613$5.22$1,905
Official, all off-peak$2.86$1,044$8.58$3,132
Official, all peak$5.72$2,088$17.16$6,263
Official, uniform traffic (21% peak)$3.46$1,261$10.37$3,784
Cheapest third party (DigitalOcean)$0.85$309$10.44$3,811
Most expensive third party (Azure)$2.66$971$22.93$8,369

Three conclusions fall out of that table.

For V4 Flash, the official API lost on price. A workload that runs around the clock pays $3.46 a day on DeepSeek and $0.85 a day on the cheapest host. That is a 4x difference on the same model, and it holds even against DeepSeek's best-case off-peak rate (3.4x). The old official price, $1.68 a day, is now available from six different hosts, so a buyer who liked the pre-August deal can still have it, just not from DeepSeek.

For V4 Pro, it is a coin flip that depends on your schedule. DeepSeek off-peak is the cheapest route at $8.58 a day. Uniform traffic costs $10.37 on DeepSeek and $10.44 on DigitalOcean, a rounding error apart. A workload concentrated in peak hours should leave the official API entirely: $17.16 a day versus $10.44.

The cheapest host is not automatically the right host. DigitalOcean's V4 Pro endpoint showed 33 tokens per second and 92 percent uptime on the day these prices were pulled, against 69 tokens per second and 99.6 percent for CoreWeave at $1.15 and $2.55. Promotional discounts (StreamLake, GMI Cloud) can end. Some hosts serve quantized weights. Price is one input to the routing decision, and the reasons it varies so much between hosts are covered in Why token prices differ.

Which DeepSeek model and host to use

Use V4 Flash when the task is retrieval, extraction, classification, or routine agent steps. It is a third of Pro's price on the official API and roughly a tenth through the cheapest hosts. Buy it from a third-party host unless you have a compliance reason to stay with DeepSeek.

Use V4 Pro when the task needs the frontier-class reasoning. Buy it from the official API if you can shift traffic off-peak, and from a third-party host if you cannot. Watch the cache-read rate: the official API charges $0.022 to $0.044 per million cached tokens, the cheapest of any host, and a workload with a 60 percent cache hit rate narrows the gap with third parties considerably.

Run the blended number, not the list price. Input share, output share, cache rate, and hour of day each move the answer. A model-by-model view of the same calculation across DeepSeek, GLM, Qwen, and Kimi is in the LLM API pricing comparison.

Older DeepSeek models: V3.2 and the retired names

DeepSeek's official API no longer lists V3.2. When V4 launched on April 24, 2026, the legacy model names deepseek-chat and deepseek-reasoner were pointed at V4 Flash, and both names were retired on July 24. The V3.2 rates of $0.28 input and $0.42 output are historical for the official API. Third-party hosts still serve V3.2 weights, and for buyers with V3.2 pinned in production that market is now the only one. The provider table is in DeepSeek V3.2 API pricing.

Model-specific pages, kept current as prices move:

Where DeepSeek pricing goes from here

The August 16 increase is the clearest example yet of how posted-price inference works. One seller changed its number, gave three days' notice on the specifics, and every buyer on the official API absorbed it. There was no market price to check the move against and no contract to hold the old price.

What happened next is the more interesting part. Seventeen other hosts serve the same weights, most of them held the old price, and the official API went from the cheapest V4 Flash host to one of the most expensive within a week. That is a competitive supply side doing its job. What it still lacks is a single price that reflects all of it: a number a buyer can quote, budget against, and lock. The Mercatus Token Index tracks where V4 Flash and V4 Pro actually trade across hosts, and How a token exchange works explains how that number gets set by trades rather than by a pricing page. For the broader picture on why Chinese open-weight models are the ones being repriced, see US vs China token prices.

Frequently asked questions

How much does the DeepSeek API cost?
As of September 2026, the official DeepSeek API charges $0.22 to $0.44 per million input tokens and $0.66 to $1.32 per million output tokens for V4 Flash, and $0.66 to $1.32 input and $1.98 to $3.96 output for V4 Pro. The lower figure applies off-peak, the higher at peak. Third-party hosts list V4 Flash from $0.0679 input and $0.168 output, and V4 Pro from $0.87 input and $1.74 output.

Did DeepSeek raise its API prices?
Yes. On August 16, 2026 DeepSeek replaced its flat rates with peak and off-peak pricing. Output tokens rose 128 to 136 percent off-peak and 355 to 371 percent at peak. Cached input tokens rose up to 12x at peak on V4 Pro. The figures were published on August 13, three days before they took effect.

What are DeepSeek's peak hours?
01:00 to 04:00 and 06:00 to 10:00 UTC, Monday through Friday. All other hours, including weekends, are off-peak and bill at half the peak rate.

Is the official DeepSeek API the cheapest way to run V4?
Not for V4 Flash. On September 2, 2026, sixteen third-party hosts listed V4 Flash below DeepSeek's off-peak rate. For V4 Pro, the official API is cheapest off-peak on input and cache reads; at peak, seven hosts undercut it on input and all sixteen on output.

Is DeepSeek V3.2 still available on the official API?
No. The deepseek-chat and deepseek-reasoner names were retired on July 24, 2026, and the official pricing page lists only V4 models. V3.2 is available through third-party hosts.

Why is the same DeepSeek model priced so differently across hosts?
Hosts run different hardware, batch sizes, quantization, and margins, and they set prices independently. DeepSeek's August increase widened the spread because most hosts did not follow it. The full breakdown is in Why token prices differ.

Methodology

Official DeepSeek rates, peak-hour definitions, and model specifications are from the DeepSeek API documentation pricing page and change log as of September 2, 2026. Pre-August 16 rates are from DeepSeek's prior pricing page and contemporaneous coverage of the August 13 announcement. Third-party host prices, throughput, and uptime figures are from OpenRouter's model pages for DeepSeek V4 Flash and V4 Pro on September 2, 2026, and reflect list prices on that date including any promotional discounts then in effect. Workload calculations assume uncached input; the uniform-traffic case weights peak hours at 35 of 168 weekly hours. Prices change; the model-specific pages linked above are updated as they do.