Kimi K3 costs $3.00 per million input tokens and $15.00 per million output tokens on Moonshot's official API, with cache hits at $0.30. Across the eleven providers hosting it, prices run from $2.80 to $3.45 input, a 1.23x spread, the tightest we have documented for any open-weight model.
That tightness is itself the story. GLM-5.2's hosting market swung from a 30x spread to 3.5x inside a month; DeepSeek repriced by up to 12x with ten days' notice. K3, six weeks after release, prices almost identically everywhere. This page has the full rates, what a real workload pays, and how the most expensive open-weight flagship stacks against the cheap ones.
What Kimi K3 is, briefly
Released July 16, 2026, with open weights following on July 27 under a modified MIT license, Kimi K3 is Moonshot AI's flagship: a 1M-token context window, vision, tool calling, and prompt caching. On public benchmark aggregates it is the strongest open-weight model we track, 97th percentile for general intelligence and 92nd for coding.
The trade-offs are price and speed: roughly 36 tokens per second with a 3.7-second time to first token, the slowest of the flagship open-weight tier. K3 is positioned as a premium reasoning model, and it is priced like one.
Kimi K3 official API pricing
| Token type | Price per 1M tokens (Moonshot official) |
|---|---|
| Input, cache miss | $3.00 |
| Input, cache hit | $0.30 |
| Output | $15.00 |
Flat pricing across the full 1M window, no long-context surcharge, and a 90 percent cache discount. At $15 per million output tokens, K3 is the most expensive open-weight model we track, roughly 3.4x GLM-5.2's official output rate and nearly 4x DeepSeek V4 Pro's peak rate.
Kimi K3 pricing by provider: all 11 hosts
| Provider | Input $/1M | Output $/1M | Cached input $/1M |
|---|---|---|---|
| Morph | $2.80 | $14.00 | $0.29 |
| DeepInfra | $2.85 | $14.25 | — |
| Parasail | $3.00 | $15.00 | $0.30 |
| Phala | $3.00 | $15.00 | $0.30 |
| Together | $3.00 | $15.00 | $0.30 |
| Chutes | $3.00 | $15.00 | $0.30 |
| BaseTen | $3.00 | $15.00 | $0.30 |
| Moonshot AI (official) | $3.00 | $15.00 | $0.30 |
| OpenRouter (routed) | $3.00 | $15.00 | $0.30 |
| Fireworks | $3.30 | $16.50 | $0.33 |
| Alibaba | $3.45 | $17.25 | $0.345 |
Seven of eleven hosts price at exactly official list, two undercut it by under 7 percent, and two charge a premium of 10 to 15 percent. Total spread: 1.23x.
Why K3's market is so tight
Three models, three market structures, in the order we documented them:
| GLM-5.2 (at launch) | DeepSeek V4 | Kimi K3 | |
|---|---|---|---|
| Cross-host spread | 30x, then 3.5x | ~3x | 1.23x |
| Cheapest seller | Third parties | Official off-peak, third party at peak | Third party, barely |
| Price behavior | 95% collapse, 10x snapback | 12x increase by announcement | Down 6.7% in 90 days |
K3's discipline likely comes down to economics rather than etiquette. A premium model with 97th-percentile quality gives hosts little reason to discount: demand is quality-driven rather than price-driven, and at $15 output the margin per token is worth protecting. There is no promotional race to the bottom because the buyers K3 attracts are not shopping on price. GLM at launch was the opposite: a commodity-priced model where hosts competed for volume with loss-leader rates that later vanished overnight.
The practical upside of a tight spread: provider choice barely matters on cost, so it can be made entirely on latency, region, and reliability. The caution from the broader spread analysis still applies, six weeks is early, and every market we have tracked has moved.
What a real workload pays
A production assistant doing 10M input and 1M output tokens per day:
| Provider | Daily | Monthly | Annual |
|---|---|---|---|
| Morph (cheapest, $2.80/$14.00) | $42.00 | $1,260 | $15,330 |
| Official Moonshot ($3.00/$15.00) | $45.00 | $1,350 | $16,425 |
| Alibaba (most expensive, $3.45/$17.25) | $51.75 | $1,552.50 | $18,889 |
The full-spread difference on this workload is about $3,500 a year, real money, but a rounding error next to the same workload's cost differences across models. On DeepSeek V4 Pro off-peak, this workload runs about $3,132 a year; on K3 it runs $15,000 to $19,000. The model choice is a 5x decision; the K3 provider choice is a 1.2x decision.
With a 60 percent cache hit rate on input, the official API drops to $28.80 a day ($10,512 a year), the 90 percent cache discount doing most of the work. For cache-heavy agent workloads, that discount matters more than anything in the provider table.
How to choose a K3 provider
- Cost: barely a decision; Morph and DeepInfra shave under 7 percent. Verify precision and rate limits as always, but the spread doesn't reward heavy shopping.
- Latency-sensitive workloads: K3 is slow everywhere (~36 tok/s); test hosts on speed rather than price, since that's where they actually differ.
- Cache-heavy workloads: the 90 percent cache discount is the biggest lever on the page; structure prompts for reuse before optimizing anything else.
- Cost-sensitive workloads generally: the honest answer is a different model. K3 buys top-tier reasoning; if the workload doesn't need it, DeepSeek V4 or GLM-5.2 deliver tokens at a fifth of the price.
- Budget certainty: August demonstrated that posted token prices move by announcement in both directions; locking a price is what forward contracts are for.
Frequently asked questions
How much does Kimi K3 cost per million tokens?
$3.00 input and $15.00 output on the official Moonshot API, with cached input at $0.30. Third-party hosts range from $2.80 to $3.45 input across 11 providers.
Who is the cheapest Kimi K3 provider?
Morph at $2.80 input and $14.00 output, about 7 percent below official list. The spread across all hosts is only 1.23x, the tightest of any open-weight model we track.
Is Kimi K3 open weight?
Yes. Moonshot released the weights on July 27, 2026 under a modified MIT license, eleven days after the API launch.
Why is Kimi K3 so expensive?
It is priced as a premium reasoning model: the strongest open-weight benchmarks we track (97th percentile intelligence, 92nd coding), a 1M context window with no surcharge, and demand that is quality-driven rather than price-driven. Hosts have little incentive to discount it.
Is Kimi K3 worth it over DeepSeek V4 or GLM-5.2?
On cost alone, no: the same workload runs roughly 5x more on K3. The premium buys top-tier reasoning and coding quality. Benchmark your actual task; if the cheaper models pass, they win on economics by a wide margin.
Does Kimi K3 charge more for long context?
No. Flat pricing applies across the full 1M-token window on the official API.
Methodology
Provider pricing compiled from published rate cards and public pricing trackers (sourced from OpenRouter and Helicone data), August 27, 2026. Benchmark percentiles from public aggregates. Workload examples use stated token volumes and cache assumptions; recompute with your own mix. Cross-model comparisons use the corresponding Mercatus pricing pages. Last verified: 2026-08-27.
