The AI Token Economy: Demand Is Decelerating While the Buildout Races Ahead
A deterministic supply-and-demand model of AI inference, anchored to what the labs actually disclose
Two things are happening at once in the AI token economy, and they point in opposite directions. Demand for AI model output, measured in tokens processed, is still growing at triple-digit rates, but the rate is decelerating quarter by quarter. Meanwhile the compute being built to serve that demand is starting to run ahead of it. Our deterministic tokenomics model, rebuilt this week from the latest company disclosures, puts hard numbers on both.
The model is bottom-up. It starts from the accelerator fleet the hyperscalers have actually deployed, converts that into a physics ceiling on how many tokens the hardware could produce, applies a realized-efficiency factor calibrated to disclosed usage, and derives token supply, inference revenue, and the balance between the two. Every disclosed number, from Google, OpenAI, and Microsoft, is used only to check the model, never to drive it. As a validation point, each disclosed figure lands at roughly 0.3 to 1.9 percent of the model's economy-wide total, which is the right scale for a single company's slice of all inference.
Demand is still exploding, but the curve is bending
The clearest new signal came from Google. On its July 22 earnings call, Alphabet said its model APIs are now processing about 22 billion tokens per minute, up from 16 billion the prior quarter. OpenAI disclosed 15 billion per minute in March, up from 6 billion at its October developer day. These are company-stated figures, and they are the cleanest read the model gets on real demand.
Fed those points, the model's fitted growth path decelerates sharply. Year-over-year token-demand growth runs near 385 percent in 2026, then roughly halves each year after: about 126 percent in 2027, 52 percent in 2028, 24 percent in 2029, and 12 percent by 2030. This is not a crash. It is the normal shape of a new technology moving from a near-zero base into the steep part of adoption and then maturing. The important part is that the deceleration is already visible in the disclosed data, not just assumed. Google's own sequence — 7 to 10 to 16 to 19 to 22 billion per minute — shows sequential growth rates of 43%, 60%, 19%, and 16%. The step down from Q1's 60% to just 16% in the latest quarter is the deceleration in one number.
The buildout is starting to outrun the demand
Here is the tension. The model tracks how many tokens the deployed fleet could serve and compares it to how many the demand path needs at forecast prices. That ratio, which we call supply cover, sits near 1.1 times in 2026 and rises to about 1.3 times by 2027 in the base case. A number above one means the hardware can produce more tokens than paying demand requires. The fleet is being built slightly ahead of the revenue that justifies it.
Modeled inference revenue still grows steeply in absolute terms. In the base case it runs about 44 billion dollars annualized in 2025, 240 billion in 2026, and 778 billion in 2027. The scenario range is wide because it turns on price per token, which falls as efficiency improves, and on how fast the fleet is utilized. The chart below shows the three cases. Note that the bear case, which assumes prices hold up, shows higher near-term revenue but the most oversupply, while the bull case has the lowest prices and the highest volumes.
Why this matters right now
The market repriced exactly this tension this week. Alphabet raised its 2026 capital expenditure guidance to a range of 195 to 205 billion dollars, up from 180 to 190 billion, and the stock fell about 7 percent despite beating on revenue and 82 percent cloud growth. Tesla fell sharply on negative free cash flow tied to AI investment. Investors are no longer just cheering the spending. They are asking when it pays back.
The model frames that question in one number: supply cover crossing above one. If it keeps climbing, the industry is provisioning compute faster than token revenue is arriving, which pressures returns on all that capex. If demand growth surprises to the upside, or efficiency gains stall and hold prices up, the gap closes. The disclosed token-per-minute figures are the fastest public gauge of which way it breaks.
Efficiency is the quiet variable
Underneath the revenue path is realized token efficiency, the share of the fleet's physics ceiling that is actually turned into useful tokens. The model calibrates it at roughly 0.5 percent in early 2025 rising to about 1.2 percent by mid-2026, roughly doubling each year. That sounds tiny, and it is, because the theoretical ceiling assumes every chip runs flat out on nothing but token generation. The gap between that ceiling and reality, driven by training versus serving allocation, utilization, and batching, is where most of the uncertainty lives. Non-serving use, mostly training and internal workloads, stays the majority of allocation through 2027 rather than flipping to inference.
What to watch
- Token-per-minute disclosures. Google and OpenAI both report these on a roughly quarterly cadence. The direction of the deceleration is the single cleanest demand signal, and it feeds the model directly.
- Supply cover. Whether the modeled buildout keeps running ahead of demand, or demand catches up. Hyperscaler capex guidance is the input that moves it.
- Price per token. Efficiency gains that lower cost per token expand volume but compress the revenue per unit. The balance between those two decides whether the bull or bear revenue path wins.
A note on confidence. This is a scenario model, not a validated forecast. The disclosed anchors bound it loosely, at roughly 2 percent of the total, so the model predicts the large majority of activity that no one discloses from hardware physics and calibration. Treat the levels as directional and the structure, decelerating demand against a buildout that is running ahead, as the durable read.
More research at bargo.ai/research.