On August 6, 2026, DeepSeek told developers a significant price increase was coming. On August 16, it landed: cache-hit input on V4 Pro rose 12x at peak hours, output rose 4.6x, and flat-rate billing was replaced by peak and off-peak windows. Ten days from announcement to effect, and for buyers on the API there was no instrument to lock the old price. You could read the announcement, and you could watch the date approach.
That sequence is the clearest case study yet for why token forward contracts exist. This piece walks through what the repricing did to a real workload's economics, what a locked price would have changed, in dollars, and how locking a token price actually works now that it's possible.
What happened, in multipliers
The verified changes, old flat rate to new peak rate, from DeepSeek's official schedule:
| Token type (V4 Pro) | Before Aug 16 | Peak after | Multiplier | Off-peak after | Multiplier |
|---|---|---|---|---|---|
| Input, cache hit | $0.003625 | $0.044 | 12.1x | $0.022 | 6.1x |
| Input, cache miss | $0.435 | $1.32 | 3.0x | $0.66 | 1.5x |
| Output | $0.87 | $3.96 | 4.6x | $1.98 | 2.3x |
Cache-heavy workloads, the agents and RAG systems that DeepSeek's famous 99 percent cache discount attracted, took the hardest hit. The full new schedule is on the V4 pricing page and the V4 Pro deep dive; the scheduling angle is covered in peak and off-peak hours, and every DeepSeek model and host is on DeepSeek API pricing.
And this was not an isolated event. The same month, GLM-5.2's cheapest host repriced 10x overnight as its discount tier vanished. Two flagship open-weight models, two double-digit repricings, one August.
The buyer's position: watch the page
A posted price is a number one seller chooses, and it can move by announcement, in either direction, at any time. That cuts both ways: DeepSeek cut prices repeatedly through 2025 and 2026 before raising them, and buyers enjoyed every cut without signing anything.
The problem is asymmetry of planning. A budget is a commitment; a posted price is not. Any team that budgeted a year of inference at August 15 rates was wrong by 64 to 230 percent on August 16, and there was nothing on the official API to do about it: posted-price APIs sell spot only, with no term structure. In the ten-day window, the available "hedge" was topping up your balance, which locks nothing.
The worked example: locked vs unlocked
Illustrative numbers, stated assumptions. Take a production workload of 10M input and 1M output tokens per day on V4 Pro, which cost $5.22 per day, about $1,905 per year, under flat-rate pricing.
Unhedged, after August 16: the same workload costs $8.58 per day ($3,132 per year) if it runs entirely off-peak, and $17.16 per day ($6,263 per year) at peak. A US-business-hours workload lands near the off-peak figure, up 64 percent; an Asia-daytime workload lands near the peak figure, up 229 percent.
Hedged, hypothetically: suppose the buyer had locked twelve months of that volume with a forward contract priced at a 5 percent premium to the pre-increase rate, about $2,000 for the year, settling in delivered tokens each month. The repricing arrives, and their cost does not move.
| Position | Annual cost | vs. flat-rate era |
|---|---|---|
| Locked forward (illustrative, +5% premium) | ~$2,000 | +5% |
| Unhedged, off-peak workload | $3,132 | +64% |
| Unhedged, peak-heavy workload | $6,263 | +229% |
The saving is $1,100 to $4,300 on a single modest workload, and the number scales linearly with volume. The honest other side: if prices had fallen instead, the way GLM's spot floor fell 95 percent before snapping back, the locked buyer pays the premium and misses the ride down. That is what a hedge is: you pay a known small cost to remove an unknown large one. Whether it's worth it is a standard finance question, and the point of a forward market is that the question can finally be asked with real numbers.
Why the seller signs the same contract
A forward needs two sides, and the seller's motivation is the mirror image. A GPU operator's costs are almost entirely fixed: depreciation runs whether the fleet is utilized or not, and idle hours are pure sunk cost. Selling capacity forward converts uncertain future utilization into contracted revenue at a known price, which also changes what lenders will finance. The buyer buys certainty about costs; the seller buys certainty about revenue; the trade exists because both kinds of certainty are worth a premium to the party that lacks them.
How locking a token price works now
Until this year, none of this was possible for tokens outside bespoke enterprise contracts. As of August 2026, physically settled token forwards trade on the Mercatus Forward Market: a price fixed at trade time for delivery in a future month, settled in actual usable tokens from a capacity-verified seller, not a cash difference against someone's posted rate. The first contracts are live, with GLM-5.2 the first model to trade. The mechanics, order book, delivery, and what the emerging forward curve signals, are in How a Token Exchange Works.
The direction of travel is bigger than one venue: the CFTC opened a request for comment on compute derivatives this month, and CME plans GPU rental futures for October. The full picture is in the CFTC analysis. Repricings like DeepSeek's are why that entire stack is being built.
Frequently asked questions
Can you lock in LLM API prices?
On posted-price APIs, no; they sell spot only. As of 2026, token forward contracts allow locking a price for future monthly delivery, physically settled in tokens. Enterprise agreements can also fix prices bilaterally, at negotiation cost and single-vendor lock-in.
How much did DeepSeek raise prices?
Effective August 16, 2026: up to 12x on cache-hit input at peak, 3x on cache-miss input, 4.6x on output, with off-peak rates at half of peak. Announced ten days earlier with no rates disclosed at announcement.
What is a token forward contract?
An agreement to buy a stated volume of inference tokens for delivery in a future month at a price fixed today. Physically settled versions deliver actual tokens from capacity-verified sellers at expiry.
Is hedging worth it if prices usually fall?
That is the trade. Open-weight token prices trended down through 2025-2026, then two flagship models repriced upward by double digits in one month. A hedge costs a known premium and removes exposure in both directions on the volume you lock; most teams hedge the predictable base load and ride spot for the rest.
Who takes the other side of a token forward?
Operators with GPU capacity: their costs are fixed, idle hours are sunk, and contracted forward revenue at a known price is worth more to them, and to their lenders, than uncertain spot utilization.
Methodology
DeepSeek rate changes verified against the official API documentation, effective August 16, 2026. Workload examples use stated token volumes; the forward premium in the worked example is illustrative and labeled as an assumption, not market data. GLM-5.2 repricing as documented on the Mercatus GLM-5.2 pricing page, August 2026. Forward market details reflect publicly announced Mercatus contracts. Last verified: 2026-08-25.
