Jensen Huang: Closed Models Are Cheaper, Open Models Are About Control
DeepSeek's $1.5B Flash ARR does not add up. The only disclosed number is $205M theoretical max. NVDA's mix still favors closed-model inference spend.
Closed is cheaper, open is about control
Nvidia CEO Jensen Huang flipped the open-source narrative last week in South Korea. Speaking to Bloomberg TV, he said closed models are often cheaper than open models once you count the full cost of hosting, fine-tuning, guardrails, evals and safety. The reason to use open models, he said, is control for specialized use cases, sovereignty and proprietary data. The interview was summarized by Benzinga and Bloomberg.
That matters for NVDA because Nvidia monetizes both sides. Closed-model pricing power funds hyperscaler capex, open-weight adoption expands the number of enterprises buying GPUs, HBM, networking and power. SemiVision framed it as AI's next war is about who controls models, not who has the best model.
DeepSeek $1.5B ARR does not check out
The claim circulating on X is that DeepSeek is doing $1.5B ARR from Flash tokens. There is no primary disclosure to support it. The post that sparked it was explicitly asking to sanity-check the math, not reporting company revenue.
What is disclosed:
- DeepSeek has never disclosed actual revenue, per CrossingRiver's synthesis of Chinese LLM economics
- The only company-disclosed figure is a theoretical max: $562k daily revenue vs $87k daily cost at H800 $2 per hour rental, implying 545% cost-profit ratio or about 84.5% margin, which annualizes to about $205M per year if all usage were billed at R1 prices. The company noted actual revenue is substantially lower due to free web and app usage and off-peak discounts, via Yahoo Finance and Computerworld
Pricing context: V4-Flash is listed at RMB 1 per million input tokens (about $0.14) and RMB 2 per million output tokens in recent coverage, while V4-Pro is RMB 3 and 6. That is roughly 1/100th of frontier closed pricing on output. At that price, $1.5B ARR would require about 5.4 quadrillion output tokens (about 8.6 quadrillion blended). Even under optimistic throughput assumptions shared in this GPU math thread, hitting $8B ARR would need about 78k to 98k GPUs at full utilization. $1.5B would still need on the order of 15k to 17k dedicated inference GPUs, far beyond any disclosed DeepSeek fleet.
Our prior work found the same gap: Chinese AI startups are beating Baidu, not OpenAI, and the $1.5B Flash ARR is a misread. Verified annualized revenue across Chinese private labs was closer to $2.6B combined versus OpenAI alone at $25B, and DeepSeek's own V4 Flash 0731 post-training jump was met by an 80% price cut from OpenAI the day before.
NVDA hyperscaler mix and closed-model inference demand
NVDA FQ1'27 ended Apr 30: revenue $81.6B, up 19.8% QoQ and 85.2% YoY, gross margin 74.9%, operating margin 65.6%, free cash flow $48.6B at 59.5% margin, net cash $68.2B. Stock $219.52 as of Aug 6, trading at 33.8x trailing and 17.1x forward earnings.
Inference economics today still favor closed-model spend. Blended list prices as of Aug 6:
Closed models are about 6x to 14x more expensive per token than open-source at list price (Claude $6.00 vs open-source $0.44), but enterprises pay for reliability, safety and lack of self-hosting burden. That is Jensen's TCO argument. OpenAI cut its blended price 19.7% in the last 30 days to $2.81, and Claude's blended price fell to $6.00, narrowing the premium versus a month ago. See live GPU rental prices for the hosting cost side of that TCO.
The observed GPU rental market as of Aug 5 is Balanced at 50.6 on the Compute Tightness Index, up 3.0 points in 30 days and tightening. H200 is Tight at 65.9, B200 Balanced at 52.4, H100 Loose at 43.2. That is a balanced market, not an acute shortage.
The enterprise adoption picture is nuanced. Morgan Stanley's note that 60% of enterprises use open-weights, covered in 60% of enterprises use open-weights. That does not mean closed is losing, does not mean closed is losing. Production data shows closed still wins spend, open wins volume. Frontier reasoning stays on closed, routine work shifts to cheap open. Both expand compute demand, which is why Rubin economics still print 78% margin even at $53 per GB HBM.
What to watch
- NVDA Q2 earnings for data center mix commentary on inference vs training and sovereign AI
- OpenAI and Anthropic pricing moves after DeepSeek V4 Flash, the price war is accelerating
- Actual DeepSeek usage disclosures, if any, versus theoretical max math
- Hyperscaler capex guides for H2, which remain the best read on whether closed-model inference demand is accelerating
More research at bargo.ai/research.