HomeBlogKimi K3 API Pricing: $3 Input, $15 Output, 11 Providers
GeneralAug 27, 20264 min read

Kimi K3 API Pricing: $3 Input, $15 Output, 11 Providers

Kimi K3 costs $3 per million input tokens and $15 output on the official API. All 11 providers, the 1.2x spread, cache math, and how it compares to DeepSeek V4.

M

Mercatus Compute

Author

Kimi K3 API Pricing: $3 Input, $15 Output, 11 Providers

Kimi K3 costs $3.00 per million input tokens and $15.00 per million output tokens on Moonshot's official API, with cache hits at $0.30. Across the eleven providers hosting it, prices run from $2.80 to $3.45 input, a 1.23x spread, the tightest we have documented for any open-weight model.

That tightness is itself the story. GLM-5.2's hosting market swung from a 30x spread to 3.5x inside a month; DeepSeek repriced by up to 12x with ten days' notice. K3, six weeks after release, prices almost identically everywhere. This page has the full rates, what a real workload pays, and how the most expensive open-weight flagship stacks against the cheap ones.

What Kimi K3 is, briefly

Released July 16, 2026, with open weights following on July 27 under a modified MIT license, Kimi K3 is Moonshot AI's flagship: a 1M-token context window, vision, tool calling, and prompt caching. On public benchmark aggregates it is the strongest open-weight model we track, 97th percentile for general intelligence and 92nd for coding.

The trade-offs are price and speed: roughly 36 tokens per second with a 3.7-second time to first token, the slowest of the flagship open-weight tier. K3 is positioned as a premium reasoning model, and it is priced like one.

Kimi K3 official API pricing

Token typePrice per 1M tokens (Moonshot official)
Input, cache miss$3.00
Input, cache hit$0.30
Output$15.00

Flat pricing across the full 1M window, no long-context surcharge, and a 90 percent cache discount. At $15 per million output tokens, K3 is the most expensive open-weight model we track, roughly 3.4x GLM-5.2's official output rate and nearly 4x DeepSeek V4 Pro's peak rate.

Kimi K3 pricing by provider: all 11 hosts

ProviderInput $/1MOutput $/1MCached input $/1M
Morph$2.80$14.00$0.29
DeepInfra$2.85$14.25
Parasail$3.00$15.00$0.30
Phala$3.00$15.00$0.30
Together$3.00$15.00$0.30
Chutes$3.00$15.00$0.30
BaseTen$3.00$15.00$0.30
Moonshot AI (official)$3.00$15.00$0.30
OpenRouter (routed)$3.00$15.00$0.30
Fireworks$3.30$16.50$0.33
Alibaba$3.45$17.25$0.345

Seven of eleven hosts price at exactly official list, two undercut it by under 7 percent, and two charge a premium of 10 to 15 percent. Total spread: 1.23x.

Why K3's market is so tight

Three models, three market structures, in the order we documented them:

GLM-5.2 (at launch)DeepSeek V4Kimi K3
Cross-host spread30x, then 3.5x~3x1.23x
Cheapest sellerThird partiesOfficial off-peak, third party at peakThird party, barely
Price behavior95% collapse, 10x snapback12x increase by announcementDown 6.7% in 90 days

K3's discipline likely comes down to economics rather than etiquette. A premium model with 97th-percentile quality gives hosts little reason to discount: demand is quality-driven rather than price-driven, and at $15 output the margin per token is worth protecting. There is no promotional race to the bottom because the buyers K3 attracts are not shopping on price. GLM at launch was the opposite: a commodity-priced model where hosts competed for volume with loss-leader rates that later vanished overnight.

The practical upside of a tight spread: provider choice barely matters on cost, so it can be made entirely on latency, region, and reliability. The caution from the broader spread analysis still applies, six weeks is early, and every market we have tracked has moved.

What a real workload pays

A production assistant doing 10M input and 1M output tokens per day:

ProviderDailyMonthlyAnnual
Morph (cheapest, $2.80/$14.00)$42.00$1,260$15,330
Official Moonshot ($3.00/$15.00)$45.00$1,350$16,425
Alibaba (most expensive, $3.45/$17.25)$51.75$1,552.50$18,889

The full-spread difference on this workload is about $3,500 a year, real money, but a rounding error next to the same workload's cost differences across models. On DeepSeek V4 Pro off-peak, this workload runs about $3,132 a year; on K3 it runs $15,000 to $19,000. The model choice is a 5x decision; the K3 provider choice is a 1.2x decision.

With a 60 percent cache hit rate on input, the official API drops to $28.80 a day ($10,512 a year), the 90 percent cache discount doing most of the work. For cache-heavy agent workloads, that discount matters more than anything in the provider table.

How to choose a K3 provider

  • Cost: barely a decision; Morph and DeepInfra shave under 7 percent. Verify precision and rate limits as always, but the spread doesn't reward heavy shopping.
  • Latency-sensitive workloads: K3 is slow everywhere (~36 tok/s); test hosts on speed rather than price, since that's where they actually differ.
  • Cache-heavy workloads: the 90 percent cache discount is the biggest lever on the page; structure prompts for reuse before optimizing anything else.
  • Cost-sensitive workloads generally: the honest answer is a different model. K3 buys top-tier reasoning; if the workload doesn't need it, DeepSeek V4 or GLM-5.2 deliver tokens at a fifth of the price.
  • Budget certainty: August demonstrated that posted token prices move by announcement in both directions; locking a price is what forward contracts are for.

Frequently asked questions

How much does Kimi K3 cost per million tokens?

$3.00 input and $15.00 output on the official Moonshot API, with cached input at $0.30. Third-party hosts range from $2.80 to $3.45 input across 11 providers.

Who is the cheapest Kimi K3 provider?

Morph at $2.80 input and $14.00 output, about 7 percent below official list. The spread across all hosts is only 1.23x, the tightest of any open-weight model we track.

Is Kimi K3 open weight?

Yes. Moonshot released the weights on July 27, 2026 under a modified MIT license, eleven days after the API launch.

Why is Kimi K3 so expensive?

It is priced as a premium reasoning model: the strongest open-weight benchmarks we track (97th percentile intelligence, 92nd coding), a 1M context window with no surcharge, and demand that is quality-driven rather than price-driven. Hosts have little incentive to discount it.

Is Kimi K3 worth it over DeepSeek V4 or GLM-5.2?

On cost alone, no: the same workload runs roughly 5x more on K3. The premium buys top-tier reasoning and coding quality. Benchmark your actual task; if the cheaper models pass, they win on economics by a wide margin.

Does Kimi K3 charge more for long context?

No. Flat pricing applies across the full 1M-token window on the official API.

Methodology

Provider pricing compiled from published rate cards and public pricing trackers (sourced from OpenRouter and Helicone data), August 27, 2026. Benchmark percentiles from public aggregates. Workload examples use stated token volumes and cache assumptions; recompute with your own mix. Cross-model comparisons use the corresponding Mercatus pricing pages. Last verified: 2026-08-27.