Bargo
NVDA

NVDA Rubin Economics: Why $8.3M per Rack Still Prints 78% Margin at $53/GB HBM

Wells Fargo says $8.34M, Bernstein says $9.1M, Morgan Stanley says 78% token margin — mapping who wins and loses

Bargo · 2026-08-05

NVDA's next rack costs almost double. Wells Fargo says Vera Rubin NVL72 is $8.34M vs $4.23M for GB300. Bernstein says even that is low — $9.1M because HBM hits $53 per GB in 2027. Yet Morgan Stanley says Rubin still prints 78% net margin selling tokens. The math only works if you look at cost per gigawatt and tokens per megawatt, not cost per rack.

The $8.3M sticker shock

Wells Fargo's Aug 4 teardown went viral on X. Per-rack total $8,341,971, up 97% or $4.23M vs GB300. Per-GW of IT power $61.9M vs $50.1M, up only 24%. That gap is the whole story. You pay double per cabinet, but you get more performance per watt.

Breakdown per rack from Wells Fargo via pequityresearch:

NVL72 BOM: GB300 vs Vera Rubin per rack ($M)
Cost per GW: still up, but less

Simple take: think of a rack like a truck. The truck costs double, but it carries 3x more packages per day. Cost per package only up 24%.

This matters because memory is not a commodity right now — HBM is sold out through 2027 and pricing power sits with the makers.

Why Bernstein says $9.1M, not $8.3M

Bernstein's note flagged by jukan05: typical VR NVL72 is $9.1M, not $8M, because media uses stale HBM price of $16.6 per GB. They expect $53 per GB in 2027 when Rubin ships volume.

Math: 72 GPUs x 192GB HBM4 = 13,824 GB per rack. At $16.6 per GB = $229k. At $53 per GB = $733k. Difference = $504k per rack just for HBM chips. Add LPDDR and SSD and you get $3.2M memory plus storage vs $2M at old prices.

HBM4 price: $16.6 to $53 per GB

Who has the HBM? SemiAnalysis: Rubin HBM from Micron = zero, 70/30 SK hynix and Samsung split. Reason: packaging yield — SK hynix 72%, Samsung 60%, Micron 47% per jukan05. Micron misses first ramp, shows up stronger on second ramp.

This is why SK Hynix printed 76% margins and the market still sold it — the pass-through gap is accounting, not demand.

Token economics: why hyperscalers still buy

Morgan Stanley's intelligence factory model from pequityresearch:

Token price reduction vs Blackwell:

Token net margin by GPU generation
Token price vs Blackwell (Blackwell=100)

Simple example: Blackwell sells 1M tokens for $10, costs $4.10, profit $5.90. Rubin sells 1M tokens for $5.30, costs $1.17, profit $4.13. But Rubin does 3x more tokens per MW, so profit per MW is higher. X bulls summarize as INFERENCE IS WILDLY PROFITABLE.

This is why hyperscalers added $1.5T backlog vs $365B capex growth in 2027 per oguzerkan — $1 capex to $5 sales over 4 to 5 years. The risk is AI data center debt hitting $7T if token volume does not 2x.

NVDA's workaround: less HBM, more system

Wukong note via tphuang: due to DRAM shortage, NVDA is offloading KV cache to SSD via BlueField-4, leaving only hot sessions in HBM. HBM4 288GB to 192GB, CPU SOCAMM 192GB to 96GB. CPX GDDR7 route shelved at GTC 2026.

SemiAnalysis June: SOCAMM per rack 55TB to 28TB, cost $7.6M to $6.8M, TCO $4.16 per hour per GPU to $3.90 per hour per jukan05.

Simple: KV cache is short term memory of a chat. Instead of keeping all chats on your desk (HBM), you move old chats to a filing cabinet (SSD). Slower, but you can ship. This helps NVDA system margin, hurts HBM demand per rack short term, and creates new demand for SSD.

Who wins and loses at $53 per GB

The fastest memory unwind of 2026 was leverage, not demand. The physics layer — memory, CoWoS, power — tightened while speculative layer deleveraged. Power is the new bottleneck.

What to watch

Sources: Wells Fargo BOM, Bernstein $9.1M and $53/GB, Morgan Stanley token margins, KV cache offload, SemiAnalysis HBM share

Get Bargo research in your inbox
One email when we publish. No spam, unsubscribe anytime.
Get Bargo research in your inbox
One email when we publish. No spam.