Bargo
AI

GPU Compute Is Tightening at the High End

H200 and B200 spot prices are rising, token demand just spiked 38% in a week, and hyperscalers are spending $725 billion this year. But the overall market is still balanced, and H100 remains loose.

Bargo

The AI compute market is tightening at the high end. GPU spot rental prices rose 6.3% week over week for the H100/H200/B200 blend, driven almost entirely by H200 and B200 demand. H200 is now in "Tight" regime with a Compute Tightness Index of 66.6, while H100 remains loose and cheap. The bigger story is token demand: it jumped 38% in a single week to 58.5 trillion tokens, confirming that cheaper inference is pulling more GPU capacity into the market.

The GPU market is splitting in two

Not all GPUs are tightening. The market is bifurcating between high-end cards that are scarce and older cards that still have slack.

GPU Spot Price (Jul 20) WoW Change CTI Regime On-Demand Unavailable
H200 $3.19/hr +7.4% 66.6 Tight 32 of 60
B200 $3.69/hr +7.3% 51.6 Balanced 42 of 38
H100 $1.67/hr +2.5% 38.7 Loose 30 of 146

H200 is the bottleneck. More than half of on-demand listings are unavailable, and its CTI has risen 10.9 points in 30 days, the fastest tightening of any GPU in the index. B200 is following the same path: spot at $3.69 is still a 49% discount to on-demand pricing at $7.18, but availability is already strained. H100 is the relief valve: 146 on-demand listings, only 30 unavailable, and spot at $1.67, a 56% discount to on-demand. If H200 and B200 keep tightening, demand will spill into H100 and its price should catch up.

The overall picture: tightening, not accelerating

The blended Compute Tightness Index sits at 47.5, up just 2.1 points from 45.4 a month ago. The regime is still "Balanced" (the 45 to 60 range). The CTI has oscillated between 43.7 and 49.0 all month with no breakout. The tightening is real but it is a grind, not a spike.

Compute Tightness Index (Last 30 Days)

Token demand just had a step change

The most important number this week is not a GPU price. It is token consumption. Weekly token volume hit 58.5 trillion on July 19, up from 42.4 trillion the week before. That is a 38% jump in seven days.

The Demand Index now sits at 248.9, up 32% in the last 30 days. Open source models are driving the surge: they account for 53.6% of all tokens consumed, up 43% month over month. The effective price per token keeps falling (now $1.97 per million tokens, down from roughly $2.47 in early July), and consumption keeps rising. This is the Jevons paradox in real time: cheaper inference means more inference, which means more GPU demand.

The capex backdrop: $725 billion in 2026

The four largest hyperscalers (Amazon, Microsoft, Alphabet, and Meta) are on track to spend a combined $725 billion on capital expenditures in 2026, according to Goldman Sachs estimates reported by Yahoo Finance. That is up 77% from $410 billion in 2025. Amazon alone is projecting $200 billion, Microsoft roughly $190 billion, Alphabet $175 to $185 billion, and Meta $115 to $135 billion.

This spending is not theoretical. It is hitting the GPU rental market now. SemiAnalysis reported in April that "all capacity coming online until August to September 2026 has already been booked." H100 one year rental contracts are up roughly 40% from their October 2025 lows. On demand capacity is sold out across all GPU types.

What to watch

The compute market is tightening at the quality end while slack remains at the commodity end. Three things will determine whether this turns into a genuine shortage:

  1. Token demand follow through. If 58 trillion tokens per week is the new baseline and not a one week spike, GPU rental pricing has further to run, especially for inference optimized cards.
  2. H100 spillover. Watch whether H100 spot prices start rising. That would signal that H200 and B200 tightness is pushing demand down the stack.
  3. Google's TPU push. Google is offering financial backstops to neo cloud GPU providers, competing directly with NVIDIA's rental ecosystem. If TPU capacity scales meaningfully, it could relieve NVIDIA specific pressure.

The SemiAnalysis thesis that GPU rental prices can keep rising because AI ROI is 5 to 10 times the compute cost is playing out in the data. The question is whether the supply response (Blackwell Ultra, Rubin, and custom ASICs) arrives before the demand curve runs away.


More research at bargo.ai/research.

Get Bargo research in your inbox
One email when we publish. No spam, unsubscribe anytime.
Get Bargo research in your inbox
One email when we publish. No spam.