Bargo
July 20, 2026

AI Inference Demand Just Surged 149% — and the GPU Market Is Tightening

Token consumption hit 58.5 trillion per week after GPT-5.6 and DeepSeek V4 launched in July. H200 rental capacity is already tight. The Jevons effect is real.

Bargo · July 20, 2026

Token demand, the best leading indicator for AI compute usage, hit 58.5 trillion tokens per week on July 18-19. That is up 149% from the 23.5 trillion baseline on May 6, and it jumped 36.7% in the last week alone. Two model launches lit the fuse: OpenAI's GPT-5.6 on July 9, and DeepSeek V4's official release in mid-July. The GPU rental market is absorbing the surge, but H200 capacity is already tight. This is not a one week spike. It is the acceleration of a trend that has been building for three months.

The chart tells the story at a glance.

Weekly Token Demand (Trillions) — GPT-5.6 Launched Jul 9

What Drove the Surge

Two major model launches landed within days of each other in July.

GPT-5.6 launched on July 9, 2026, with three tiers: Sol (the flagship), Terra (mid-range), and Luna (budget). The launch was delayed from June after a US government security review, and CNBC reported that the public release followed the completion of that review. The token spike hit the week of July 18-19, roughly 7 to 10 days after the API went live, which is consistent with developer ramp-up time.

DeepSeek V4 moved from preview to official release in mid-July, with the old endpoints (deepseek-chat, deepseek-reasoner) set to retire on July 24. DeepSeek introduced peak hour pricing at 2x the baseline during Beijing business hours, a clear signal that they are seeing real capacity constraints. This feeds directly into the open source token surge.

Claude Fable 5 launched June 9, but its token volume is only up 6% over the last 30 days. It is not driving this wave.

Who Is Consuming the Tokens

Open source models now dominate. They account for 53.6% of all tokens tracked, up from 35.9% in late May, with volume up 43% in the last 30 days. At $0.45 per million tokens, open source inference is less than one tenth the cost of Claude. The price is so low that it unlocks use cases that were not economical before.

OpenAI token volume jumped 40% in the last 30 days. Despite commanding only 7.2% of total tokens, OpenAI's share of implied spend is much larger because its blended price is $3.50 per million tokens, roughly 8x the open source rate.

Google is losing share. Its token volume is down 5% over 30 days. xAI is tiny in absolute terms (0.8% share) but growing at 289% off a small base.

The effective blended price across all providers has fallen to $1.97 per million tokens. That is down from roughly $2.48 in mid-June. The Jevons effect is at work: cheaper inference drives more total consumption, not less.

The GPU Market Response

The Compute Tightness Index (CTI) tracks how tight GPU rental capacity is by comparing spot prices to on-demand prices and counting unavailable listings. A score above 60 means the market is tight. Above 80 is acute.

Compute Tightness Index — H200 Already Tight, Blended Still Balanced

The blended CTI sits at 47.5, still in the Balanced zone. But the aggregate masks what is happening at the GPU level. H200 is already tight at 66.6, up 10.9 points in the last 30 days. Its spot discount is only 25% compared to 56% for the older H100, which means spot buyers are paying close to on-demand prices for H200. B200 is balanced at 51.6. H100 is loose at 38.7, with 146 on-demand listings acting as a release valve.

The spot pricing confirms the tiering.

GPU Spot Rental Prices (Jul 20, 2026)

H100 is the bargain bin at $1.67 per hour on spot. H200 spot at $3.19 is nearly double. B200 commands $3.69 on spot, but the on-demand spread is massive at $7.18, reflecting the premium for Blackwell architecture. AMD's MI300X sits at $1.65 spot, competitive with H100 on price but with far fewer listings (3 spot vs 39 for H100).

What This Means

The inference demand surge is real, broad, and still building. It is not a one week spike. The trend has been climbing since early May, and the July model launches accelerated it. The GPU market is absorbing the surge for now, but H200 tightness is the early warning. If token demand keeps climbing at this rate, capacity constraints will spread to B200 and eventually tighten the blended market.

The open source share shift is the most important structural change. At 53.6% of tokens and growing, open source inference is becoming the dominant consumption mode. This benefits infrastructure providers (cloud, GPU rental) more than any single model company. It also means the marginal cost of inference keeps falling, which keeps expanding the market.

The next data point to watch is whether the Jul 18-19 spike holds or fades. If token demand stays above 50 trillion per week through the end of July, it confirms a new plateau rather than a launch week spike. The H200 CTI crossing 70 would be the next escalation signal.


More research at bargo.ai/research.

Get Bargo research in your inbox
One email when we publish. No spam, unsubscribe anytime.
Get Bargo research in your inbox
One email when we publish. No spam.