Blog
Insights on AI compute markets, GPU pricing, and platform updates
DeepSeek API Pricing: Official Rates vs Third-Party Hosts
DeepSeek V4 Flash costs $0.22 to $1.32 per million tokens on the official API and as little as $0.07 through third-party hosts. Every DeepSeek model, every host, and what the August 16 price increase changed.
LLM API Pricing Comparison: 6 Models, Every Provider
Official rates, cheapest hosts, cache terms, and spreads for DeepSeek, GLM, Kimi, and Qwen, compared in one table. Updated from the Mercatus model pricing pages.
Kimi K3 API Pricing: $3 Input, $15 Output, 11 Providers
Kimi K3 costs $3 per million input tokens and $15 output on the official API. All 11 providers, the 1.2x spread, cache math, and how it compares to DeepSeek V4.
How to Lock In Token Prices: The DeepSeek 12x Case Study
DeepSeek raised token prices up to 12x with ten days' notice. What locking in with a forward contract would have looked like, and how token hedging works now.
GLM-5.2 API Pricing: The 30x Spread Just Collapsed to 3.5x
GLM-5.2 now costs $0.40 to $1.40 per million input tokens across 21 hosts. Two weeks ago the spread was 30x. What repriced, who got burned, and the new rates.
Compute Derivatives: The CFTC Just Asked How to Build Them
The CFTC is asking how compute derivatives should work, and CME plans GPU futures for October. What the RFC says, the settlement problem, and who fills the gap.
DeepSeek Peak and Off-Peak Hours: When Tokens Cost Half
DeepSeek's peak hours converted to your time zone, the savings math ($3,131 a year on one workload), and the batch jobs worth moving. Weekends are always off-peak.
DeepSeek V4 Pro API Pricing: Rates After the Increase
DeepSeek V4 Pro now costs $0.66 to $1.32 per million input tokens after the August 16 increase. New peak/off-peak rates, cache math, and what changed for buyers.
DeepSeek V4 API Pricing: New Peak and Off-Peak Rates
DeepSeek V4 Pro costs $0.66 to $1.32 per million input tokens since the August 16 increase. Before it, DeepSeek was the cheapest of 16 hosts. Here is the provider table now, and who wins at each hour.
US vs China Token Prices: A 7x Gap Fell to 4x in 30 Days
US inference costs $4.56 per blended million tokens. Chinese inference costs $1.12. A month ago the gap was 7x. TPI data on the fastest convergence in AI pricing.
DeepSeek V3.2 API Pricing: Official Rates vs the Cheapest Providers
DeepSeek V3.2 listed at $0.28 per million input tokens before DeepSeek retired it from the official API. Fifteen third-party hosts now serve it from $0.21. Full provider comparison, cache math, and blended cost.
Same Model, 3x the Price: Why LLM Token Pricing Varies So Much
The same open-weight model costs up to 3x more depending on the provider. Where the spread comes from, what a fair token price is, and how buyers should compare.
How a Token Exchange Works: Spot, Forward, and Price Discovery for AI Inference
How a token exchange works for AI inference: spot markets, forward contracts, order books, and public price discovery. What it changes for compute buyers locking costs and GPU operators monetizing spare capacity.
NVIDIA B200 Server Price in 2026: What an 8-GPU HGX System Costs
An 8-GPU HGX B200 server costs $400,000 to $500,000 in 2026, if you can get an allocation. What's inside the price, the 7x cloud spread, and the math vs H200.
NVIDIA H100 Server Price in 2026: What an 8-GPU HGX System Costs
An 8-GPU HGX H100 server costs $250,000 to $320,000 in 2026. What's inside the price, what resale value does to the math, and when renting beats owning.
NVIDIA H200 Server Price in 2026: What an 8-GPU HGX System Costs
An 8-GPU HGX H200 server costs $320,000 to $420,000 in 2026. Full breakdown: what's inside the price, OEM differences, cost per GPU-hour, and buy vs rent math.
NVIDIA H100 Resale Value 2026: Used Sells at $22K–$25K
What used NVIDIA H100s resell for in 2026 ($22K–$25K at 18–30 months), where to sell them, and how residual value swings your real cost of ownership.
Hidden Cloud GPU Costs in 2026: The 7 Charges Behind the Sticker $/hr
Hidden cloud GPU costs add 30 to 50 percent to a typical managed bill. The seven charges behind the gap, with 2026 rates and a formula to price your own.
B200 vs H100: When to Buy Blackwell in 2026
B200 ships in early-cohort volumes through 2026, mostly to hyperscalers. The decision framework for when Blackwell actually beats H100 on cost in 2026.
GPU ROI in 2026: Payback Period, IRR, and NPV for AI Infrastructure
How institutional buyers model the return on owning AI compute. Payback period, IRR, and NPV for a 100 H100 cluster in 2026, with current pricing, formulas, and sensitivity tables.
Financing AI Compute in 2026: Buy, Lease, Loan, or Rent
A 100 H100 cluster runs $3.3M to $4.2M fully built. Compare the five GPU financing paths in 2026 and the formula that decides which one wins.
The Cheapest GPU Cloud Providers in 2026: Where AI Compute Is Actually Lowest
Looking for the cheapest GPU cloud providers in 2026? This guide compares H100, A100, and H200 pricing across hyperscalers, specialty GPU clouds, long-tail providers, and decentralized networks — including spot, reserved, and on-demand rates.
Cloud GPU Pricing in 2026: What You Actually Pay Across All Major Providers
Cloud GPU pricing varies widely across providers. This guide breaks down H100, A100, H200, reserved pricing, hidden costs, and what buyers actually pay in 2026.
Why GPU Prices Differ 30%+ for the Same Hardware (and What It Says About the Market)
GPU prices can vary dramatically across providers for the same hardware. This guide explains why the spread exists, what drives the premium, and what it says about the AI compute market.
GPU Utilization: The Most Important Metric in AI Infrastructure (and Why Most Teams Measure It Wrong)
GPU utilization is the biggest controllable cost lever in AI infrastructure. Learn what utilization actually means, why most teams measure it wrong, and how idle GPU capacity can now be monetized in 2026.
H200 Buy vs Cloud: When Does Owning H200 Make Financial Sense? (2026)
H200 buy vs cloud comparison across utilization, fleet size, and infrastructure costs. See when owning H200 becomes cheaper than renting.
NVIDIA H200 Price 2026: $32K GPU, $370K HGX Server
NVIDIA H200 costs $32K to $40K per GPU and $320K to $420K for an 8-GPU HGX server in 2026. Compare OEM, cloud, and total ownership cost.
H100 vs H200: Is the Upgrade Worth the 25% Premium? (2026 Decision Guide)
H100 vs H200 comparison across memory, bandwidth, pricing, and inference performance. See when upgrading to H200 makes sense in 2026.
A100 vs H100: Should You Pay for Hopper or Stick with Ampere? (2026)
A100 vs H100 comparison across cost, performance, FP8 support, and real AI workloads. See when H100 justifies the premium and when A100 still makes sense in 2026.
Colocation Economics for AI Compute: From $/kW/Month to GPU Cost per Hour
Colocation pricing for AI compute ranges from $80/kW/month wholesale to $300/kW/month in premium metro facilities. Here’s how colocation economics translate into real GPU cost per hour, what drives pricing variance, and where institutional AI operators can optimize infrastructure costs.
H100 Depreciation: How Fast NVIDIA H100s Lose Value (and What It Means for TCO)
How fast do NVIDIA H100 GPUs lose value? This guide breaks down H100 depreciation rates, Blackwell’s impact on residual value, secondary market pricing, and how depreciation affects real AI infrastructure TCO in 2026.
The 3-Year TCO of Owning 100 H100 GPUs: A Full Capital Allocation Breakdown
A full breakdown of the 3-year total cost of owning 100 NVIDIA H100 GPUs, including hardware, power, colocation, networking, operations, depreciation, and utilization economics.
Buy vs Rent GPUs: The 2026 Decision Framework for AI Infrastructure
A detailed 2026 framework for deciding whether to buy or rent GPUs for AI infrastructure. Compare H100 ownership economics, reserved cloud pricing, utilization thresholds, and how capacity monetization changes the break-even point.
NVIDIA H100 Price 2026: $25K GPU, $285K HGX Server
NVIDIA H100 costs $25K to $30K per GPU and $250K to $320K for an 8-GPU HGX server in 2026. Compare OEM, cloud, and total ownership cost.
A100 vs H100 vs H200: Complete Cost & Performance Comparison for AI Workloads (2026)
A100 vs H100 vs H200 comparison across cost, performance, and real workloads. See which GPU makes sense for training, inference, and scale.
The Case for Open Price Indices in AI Compute
AI compute is becoming a commodity. Its benchmark indices should be public goods — open, transparent, and free — just like every other mature market.
No posts in this category yet.