HomeBlogHow a Token Exchange Works: Spot, Forward, and Price Discovery for AI Inference
GeneralJul 28, 202613 min read

How a Token Exchange Works: Spot, Forward, and Price Discovery for AI Inference

How a token exchange works for AI inference: spot markets, forward contracts, order books, and public price discovery. What it changes for compute buyers locking costs and GPU operators monetizing spare capacity.

M

Mercatus Compute

Author

How a Token Exchange Works: Spot, Forward, and Price Discovery for AI Inference

Every input that matters to the global economy trades on an open market. Oil has benchmark prices that print continuously. Electricity clears on day-ahead and forward markets. Freight rates are assessed and hedged daily. AI inference, which is becoming one of the largest operating costs in software, has none of this. You pay whatever price the provider posts, with no visible market price, no forward curve, and no way to lock costs beyond a bespoke enterprise contract.

That is starting to change, and the change follows a pattern markets have repeated for a century. This article explains how a token exchange works: what trades, how spot and forward markets function mechanically, what public price discovery changes for buyers and sellers, and why the structure of open weight models makes this possible now.

Why inference is priced like a commodity but sold like enterprise software

LLM inference has the properties of a commodity. The unit is standardized (tokens, priced per million), the product is fungible (the same open weight model produces the same output regardless of who serves it), and there are many producers (any operator with GPUs can serve an open model). Yet the market structure looks nothing like a commodity market.

Market featureOil, power, freightLLM inference today
Visible spot priceBenchmark prices print continuouslyEach provider posts its own list price
Forward curve12+ months visible to everyoneNone
Hedging instrumentsFutures and forwards are standard practiceBespoke enterprise contracts, if anything
Many sellers per productYes, with open participationClosed models have one seller; open weight supply is fragmented
Standardized settlementYesNone

The result is exactly what you would expect from a posted-price market: wide, persistent spreads for identical products. The same pattern already exists one layer down the stack, where the same H100 GPU rents for $1.99 to $5.00 per hour depending on provider, a 2.5x spread for identical hardware. The reasons GPU prices differ this much for the same hardware apply directly to tokens: no transparent venue where prices clear, so buyers overpay and sellers with cheap capacity struggle to reach demand.

Every open market started with posted prices

The current state of inference pricing is not unusual. It is the normal starting point for a commodity market. Oil traded for most of the 20th century at posted prices set by a handful of major producers. Buyers paid the list price because there was nowhere else for a price to form. It took the supply shocks of the 1970s to push cargoes into spot trading, and by 1983 crude oil futures were listed on the NYMEX. Within a decade, the posted price was dead and the market price was the only price that mattered.

Electricity followed the same arc when deregulation in the 1990s replaced utility tariffs with day-ahead auctions and forward markets. Ocean freight went from privately negotiated rates to published indices and forward freight agreements. In every case the sequence was identical: posted prices, then a spot market, then standardized forward instruments, then a public curve that the entire industry plans against.

AI inference is at step one of that sequence. The interesting question is not whether the sequence repeats, it is who builds the venue where it happens, and the historical answer is that the venue that prints the trusted price becomes the market.

What actually trades: the token as a unit

A token is the natural unit of account for AI inference. Models are priced per million input and output tokens, usage is metered in tokens, and capacity planning increasingly happens in tokens per second rather than GPU-hours.

The distinction that makes an exchange possible is open weight models. A closed model has exactly one seller, so there is no market to make, only a list price to accept. An open weight model can be served by any operator with the hardware to run it: GPU clouds, labs with spare training capacity, resellers, regional providers. Same model, same output, many competing sellers. That is the precondition for real price discovery, and it is why exchange activity concentrates on open weight models first.

The spot market: capacity meets demand at a market price

A spot token market works like any electronic marketplace. Sellers list available inference capacity for a given model at their asking price. Buyers take the best available price, filtered by whatever else they care about, typically latency and region. Access runs through a single API, so routing to the best qualified seller does not require integrating each provider separately.

Two things distinguish this from the aggregators most teams use today.

First, sellers set their own prices and compete openly. An aggregator resells a fixed menu of providers at posted prices plus a fee. An open spot market lets anyone with verified capacity list, which pulls in supply that posted-price channels never reach: idle fleet hours, off-peak regional capacity, spare headroom on reserved clusters. Most GPU fleets run well below full utilization, and utilization is the single biggest driver of GPU economics, so this supply is real and motivated.

Second, every trade prints. When transactions clear on an order book, the market produces a public price per model, continuously. That number does not exist today. Posted prices are not market prices, in the same way a hotel rack rate is not what rooms actually clear at.

Inside the order book

The mechanism underneath is a central limit order book, the same structure that runs equities, futures, and every serious electronic market. It is worth understanding because it explains where the price actually comes from.

For each model, the book holds two stacks of orders. On one side, asks: sellers offering a quantity of tokens at a minimum price, for example 500M tokens of a given open weight model at $0.61 per million. On the other side, bids: buyers willing to pay up to a maximum price. The highest bid and the lowest ask define the market at any moment, and the gap between them is the spread.

When a bid and an ask cross, the trade executes automatically and prints to the public tape. Orders are matched by price first, then by time of arrival, so no participant gets preferential treatment. A buyer who just wants tokens now takes the best ask and is done. A seller with idle capacity and patience can sit lower in the book and wait for demand to come to them.

Three properties fall out of this structure. The book runs 24/7, because inference demand does not keep business hours. The spread itself is information: a tight spread means a liquid, competitive market, a wide spread means uncertainty or thin supply. And because every fill prints, the market price per model is not an estimate or a survey, it is a record of what buyers actually paid.

The forward market: locking inference costs up to 12 months out

The forward market is where the exchange structure earns its place in a finance conversation. A forward contract lets a buyer lock a price for tokens delivered in any of the next 12 months, at a price set by the market rather than negotiated bilaterally. Contracts are backed by capacity-verified sellers and settle in actual usable tokens, not cash differences.

For buyers, this converts a volatile operating cost into a fixed one. If inference is a top-three line item, an unhedged position means your gross margin moves with someone else's pricing decisions. The logic mirrors the buy vs rent GPU decision: committing capital buys certainty, staying flexible costs a premium, and the right answer depends on how predictable your demand is. Forwards give you that dial at the token level, without owning hardware.

For sellers, a forward contract is booked revenue against capacity that would otherwise sit idle. A GPU owner's costs are almost entirely fixed: the hardware is bought, depreciation runs regardless of utilization, and colocation and power bills arrive monthly. Selling forward turns uncertain future utilization into contracted cash flow, which also changes the financing conversation for compute fleets, because lenders price contracted revenue very differently from speculative capacity.

Spot vs forward vs reserved contracts vs posted prices

Teams buying inference today have effectively two options: pay the posted price, or negotiate a committed contract with a single provider. An exchange adds two more. The four differ on who sets the price, how long you are locked, and what you are exposed to.

Posted price (status quo)Reserved / committed dealExchange spotExchange forward
Price set byProvider list priceBilateral negotiationMarket, continuouslyMarket, per delivery month
TermNoneTypically 1 to 3 yearsNone1 to 12 months
Price transparencyList price onlyPrivate, unbenchmarkablePublic printPublic curve
FlexibilityFull, switch anytimeLocked to one providerFull, best price per requestLocked volume, spot for overflow
CounterpartyOne providerOne providerMany competing sellersCapacity-verified sellers
Best forSmall or unpredictable usageLarge stable usage, single vendor comfortCost optimization at any scaleBudget certainty without vendor lock-in

The reserved deal is the closest thing today's market has to a hedge, and its weakness is exactly what an exchange fixes: you lock into one counterparty at a price you cannot benchmark, and the discount you negotiated is invisible to you the moment you sign. A forward contract does the same economic job with a market-set price and no single-vendor dependency.

Worked example: hedging a $360K annual inference budget

Illustrative numbers, stated assumptions. Say a software company runs 50 billion blended tokens per month on an open weight model at a spot price of $0.60 per million tokens. That is $30,000 per month, $360,000 per year, and growing.

// text
Unhedged annual cost = sum of (monthly tokens x that month's spot price)
Hedged annual cost   = (contracted tokens x forward price) + overflow bought at spot

The company locks 40B tokens per month for 12 months at a forward price of $0.63 per million, a 5 percent premium to today's spot, and buys overflow above 40B at spot.

Scenario over 12 monthsFully unhedged80% hedged at $0.63
Spot stays flat at $0.60$360,000$374,400
Spot rises 40% mid-year (avg $0.72)$432,000$388,800
Spot falls 20% (avg $0.48)$288,000$331,200

The hedge costs about $14,000 in the flat case and saves about $43,000 in the spike case. Whether that trade is worth it is a standard finance question, and that is precisely the point. Today the question cannot even be asked, because there is no forward price to evaluate. A forward curve turns "hope prices behave" into a decision with numbers attached.

The seller side: what idle capacity costs and what forwards change

The buy side gets most of the attention, but the seller economics are what make an open exchange work, because they explain why supply shows up at competitive prices.

A GPU operator's cost structure is dominated by fixed costs. Consider a single H100 on an owned cluster, using assumptions consistent with our H100 cost breakdown and cluster TCO analysis: roughly $28,000 in hardware capex, with power, colocation, networking, and operations adding 50 to 80 percent on top over a 3-year life. Call it $46,000 all-in over 3 years, or about $1.76 per GPU-hour if the card ran every hour of its life.

It never does. And because nearly all of that cost is fixed, the cost per hour the operator actually sells rises sharply as utilization falls:

// text
Effective cost per utilized GPU-hour = all-in hourly cost / utilization rate
UtilizationEffective cost per utilized GPU-hour (illustrative)
100%$1.76
85%$2.07
60%$2.93
40%$4.40

An operator running at 60 percent utilization is not earning anything on 40 percent of the hours they have already paid for. Every one of those idle hours is pure sunk cost, which is why idle capacity is the most motivated supply in any market. Selling it at nearly any price above marginal cost (power, mostly) improves the fleet's economics.

This is what the two market structures do for a seller. The spot market converts idle hours into revenue at the market price with no sales team and no enterprise contract cycle. The forward market goes further: by selling capacity months ahead at a locked price, the operator raises their utilization floor, turning a fleet that might run at 60 percent into one contractually committed at 80 percent or more. That shift moves the effective cost per sold hour down the table above, and it converts a speculative asset into one with visible, contracted cash flow, which is the difference between a hard and an easy conversation about GPU return on investment.

What public price discovery changes

A printed price per model, spot and 12 months forward, does work that no amount of provider comparison shopping can do.

For buyers, it establishes what a fair price is. When the same model prices far apart across providers, the spread is invisible unless you check every provider yourself. A market price makes overpaying obvious.

For sellers, it prices their asset. A GPU fleet's revenue potential becomes visible as a curve rather than a guess, which feeds directly into cluster investment math and buy versus expand decisions.

For the market, the forward curve becomes a signal. A curve that slopes upward says the market expects compute to stay scarce. One that slopes down says supply is coming. Every mature commodity market treats this curve as core economic information. AI compute currently makes these bets in the dark.

The market is already moving toward treating tokens as a financial primitive. In July 2026, the Wall Street Journal reported that Stripe is in talks to acquire OpenRouter, an inference routing marketplace, for roughly $10 billion. Routing is not price discovery, an aggregator picks a path while an exchange prints a price, but a payments company valuing token flow at that scale says clearly where inference spend is headed: metered, routed, and eventually priced like the commodity it is.

Who uses a token exchange

Buy side: engineering teams wanting the best executable price through one API, and finance and procurement teams who want inference to behave like other major input costs, forecastable, hedgeable, and benchmarked against a public price.

Sell side: anyone operating verified capacity. GPU clouds filling unsold inventory, labs monetizing capacity between training runs, enterprises with reserved clusters running below capacity, regional operators with cheap power but no sales channel. For an operator, the exchange is a distribution channel with no sales overhead, and the forward market is a way to book revenue ahead of delivery.

Mercatus publishes the data layer for both sides today: the Token Index benchmarks inference token costs across providers, and the GPU Index tracks the GPU-hour prices that set the floor under every token price.

Frequently asked questions

What is an AI token exchange?

A marketplace where LLM inference tokens are bought and sold at market-determined prices rather than provider list prices. Sellers list verified capacity for specific models, buyers purchase spot or forward through one API, and every trade contributes to a public price per model.

What is the difference between spot and forward token buying?

Spot means buying tokens for immediate use at the current market price. Forward means locking a price now for tokens delivered in a future month, up to 12 months out. Spot optimizes for the best price today; forward optimizes for cost certainty.

How is a token exchange different from an API aggregator?

An aggregator routes your request across a fixed set of providers at their posted prices, plus a fee. An exchange is open on the supply side, anyone with verified capacity can list at their own price, and it produces a public market price from actual trades. Aggregators optimize routing; exchanges optimize price discovery.

What is a forward curve and what does it tell you?

The set of market prices for delivery in each future month, plotted over time. An upward-sloping curve means the market expects tokens to get more expensive, typically a signal of tight compute supply. A downward slope signals expected supply growth or falling costs. Commodity markets treat this curve as core planning information for capacity and procurement decisions.

Why do token prices for the same model vary across providers?

Because there is no transparent venue where prices clear. Providers have different hardware costs, utilization rates, and margin targets, and buyers cannot easily compare or switch. The same dynamic produces a 2.5x spread in GPU-hour prices for identical hardware.

Can you hedge inference costs today?

Outside of bespoke enterprise agreements, inference has generally not been hedgeable, there has been no standard instrument and no market-set forward price. A forward market with capacity-verified sellers and token settlement is what makes hedging a standard practice rather than a negotiation.

How do sellers price forward contracts?

From their cost basis and utilization expectations. An operator knows their all-in cost per GPU-hour and what utilization they can commit to. Any forward price above that effective cost is margin, and a locked contract is often worth accepting at a discount to expected spot because contracted revenue is worth more than uncertain revenue, to the operator and to their lenders.

Who can sell on a token exchange?

Any operator with verifiable inference capacity for listed models: GPU clouds, AI labs, resellers, and enterprises with spare capacity on owned or reserved hardware. Capacity verification and delivery standards determine listing eligibility, so contracts are backed by real, tested supply.

Methodology

GPU-hour pricing figures are drawn from the Mercatus GPU Index, which tracks cloud GPU pricing across 30+ providers, refreshed continuously. Seller cost figures are illustrative and use assumptions consistent with Mercatus TCO analysis: roughly $28,000 H100 capex with 50 to 80 percent infrastructure overhead over a 3-year life. Hedging figures in the buyer example are illustrative and labeled as assumptions, not market data. The Stripe and OpenRouter figure reflects Wall Street Journal reporting from July 2026 on a deal in talks, not a completed transaction. Market mechanics described reflect the Mercatus exchange design: spot marketplace, 12-month forward contracts, central limit order book, capacity-verified sellers, and token settlement.