Chinese AI Startups Are Beating Baidu, Not OpenAI. The $1.5B Flash ARR Is a Misread
DeepSeek V4 Flash at $0.14 per million tokens makes $8.69 per GPU hour at sparse attention speed, but loses money at normal reasoning speed. Verified revenue is $500M annualized, not $1.5B. Chinese private labs combined do $2.6B versus OpenAI alone at $25B.
DeepSeek V4 Flash costs $0.14 per million input tokens. That is 9 times cheaper than GPT-5 and 21 times cheaper than Kimi K3. At 11,500 tokens per second per GPU, it makes $8.69 per GPU hour and keeps 77 percent margin against $2 per hour H800 rental. At normal reasoning speed, it loses money. This is the real story behind the Chinese startup surge.
The claim that DeepSeek does $1.5B ARR from Flash tokens does not check out. The $1.5B number is a reported raise amount, not revenue. Verified disclosures show theoretical max of $205M per year if all usage was billed, with press leaks pointing to around $500M annualized. OpenAI is at $25B ARR as of February 2026.
The pricing war in one table
| Model | Input per 1M | Output per 1M | Cache hit input | Context |
|---|---|---|---|---|
| DeepSeek V4 Flash | $0.14 | $0.28 | $0.0028 | 1M |
| DeepSeek V4 Pro | $0.435 | $0.87 | $0.0036 | 1M |
| Kimi K3 2.8T MoE | $3.00 | $15.00 | $0.30 | 1M |
| GPT-5 | $1.25 | $10.00 | $0.125 | 400K |
| GPT-5.6 Terra | $2.50 | $15.00 | — | 1M |
| Claude Sonnet 4.6 | $3.00 | $15.00 | — | 200K |
| Claude Opus 4.8 | $5.00 | $25.00 | $0.50 | 1M |
Sources: DeepSeek API Docs, OpenRouter Kimi K3, OpenAI Pricing
DeepSeek sets the floor. Kimi proves Chinese labs can charge premium when they have frontier capability. The blended market shows open source at $0.44 per 1M versus OpenAI at $3.50 and Claude at $10.00. Open source is 71.7 percent of tracked token volume, up from 59.8 percent in May. Cheaper tokens drive more usage, but revenue concentrates in premium tiers.
This dynamic was visible when DeepSeek V4 Flash 0731 landed and OpenAI had already cut Luna pricing 80 percent the day before. The inference price war is on.
Unit economics: Flash only works with sparse attention
We modeled revenue per GPU hour at three throughput levels using verified costs. H100 full cost is $2.215 per hour including capex recovery and financing, cash cost $0.736 per hour, spot rental $1.72 today. H800 rental in China is $2.00 per hour per DeepSeek disclosure.
| Model | Throughput | Rev per GPU hr | Cost $2.00 | Margin | Cost $2.215 | Margin |
|---|---|---|---|---|---|---|
| V4 Flash | 11.5k tok/s sparse | $8.69 | $2.00 | 77.0% | $2.215 | 74.5% |
| V4 Flash | 8.3k tok/s generic | $6.25 | $2.00 | 68.0% | $2.215 | 64.6% |
| V4 Flash | 2.2k tok/s Pro speed | $1.66 | $2.00 | -20.2% | $2.215 | -33.2% |
| V4 Pro | 8.3k tok/s | $19.41 | $2.00 | 89.7% | $2.215 | 88.6% |
| Kimi K3 | 8.3k tok/s | $267.8 | $2.00 | 99.3% | $2.215 | 99.2% |
Key insight: Flash is profitable only if sparse attention delivers 11.5k tokens per second. At normal reasoning speed of 2.2k, it loses money. That is why DeepSeek eliminated off peak discounts in September 2025. They were subsidizing low throughput workloads.
The leaked Liang Wenfeng playbook showed 10 month GPU payback at 85 percent margins, but that was at V3 pricing of $0.27 per 1M input. At Flash pricing, margin compresses unless V3.2 sparse attention cuts input 50 percent and output 75 percent versus V3.1.
What $1.5B ARR actually requires
At $0.21 per 1M average price for Flash:
- Tokens needed per year: 7.14 quadrillion
- GPUs needed at 100 percent utilization: 19,696 at 11.5k tok/s, 27,401 at 8.3k tok/s
- At 80 percent practical utilization: 24,600 to 34,250 GPUs
- For $8B ARR as Tex modeled: 105k to 146k GPUs
Global fleet is 35.5M GPUs active in 2026 per TCO model, 47.6 GW. 20k GPUs is 0.06 percent of global fleet, feasible for DeepSeek, but would be their entire inference fleet dedicated to Flash.
Revenue reality: $2.6B combined versus $25B alone
| Lab | Latest ARR | Valuation | P/S |
|---|---|---|---|
| OpenAI | $25B Feb 2026 | $852B post Mar 2026 | 34x |
| Anthropic | $7B Oct 2025 | $965B post May 2026 | — |
| Kimi Moonshot | $100M Mar 2026 to $300M Jun | $20B to $35B | 100x |
| Zhipu | $53M H1 25 ann to $153M H2 25 ann | $65.7B | 400x plus |
| MiniMax | $30.5M 2024, $53.4M 9M 2025 | $32.9B | 600x plus |
Zhipu: $8.4M in 2022 to $45.9M in 2024 revenue, losses $347M, R&D 835 percent of revenue. MiniMax: $0.25M in 2023 to $30.5M in 2024, gross margin negative 24.7 percent to positive 23.3 percent, net loss $512M. These companies cannot afford to match Flash pricing if they rent H100s at $4.01 on demand.
Kimi is different. Kimi K3 is a compute hog at 2.8T parameters, burning three times the tokens of peers, but charges $3 per 1M input. That is premium pricing for frontier capability, not a price war. At $200M ARR in April to $300M in June, with $20B to $35B valuation, they are betting 1M context and multimodal reasoning justifies US level pricing.
Why Tex is half right
@teortaxesTex wrote: Chinese AI has exceeded the American pattern of rich internet companies sucking. Small startups DeepSeek Kimi Zhipu are jostling for supremacy. Baidu Tencent Alibaba ByteDance get dunked on.
He is right about incumbents. Baidu Ernie, Tencent Yuanbao at 73M MAU, Alibaba Qwen have not converted distribution into leadership. ByteDance Doubao leads with 170M MAU despite not being best model, per platform integration. In the US, OpenAI Anthropic Google have deeper moats: cloud distribution, 920M ChatGPT users, enterprise contracts, and $16B compute spend.
He is wrong about exceeding the American pattern on revenue. Combined Chinese private labs at $2.6B run rate are still behind OpenAI alone at $25B. The valuation multiples tell the story: 100x to 600x forward P/S in China versus 34x for OpenAI. The market is pricing Chinese startups as if they will capture US level economics, but their pricing strategy makes that hard.
The funding structure explains why they can survive: National IC Big Fund may lead DeepSeek $7.3B round at $51.5B, first ever foundation model bet, elevating LLMs to semiconductor level priority. Industrial capital from handset supply chain, Tencent, China Mobile for distribution. This is state and industrial capital, not just VC.
For more on open versus closed economics, see 60 percent of enterprises use open weights and Qwen 3.8 Max reaction.
What to watch
- DeepSeek V4.5 and Kimi K3 open weight release: will sparse attention throughput hold at 11.5k tok/s in production
- H100 spot at $1.72 today, 57 percent discount to on demand $4.01: if it tightens, Flash margin compresses for anyone renting
- Zhipu and MiniMax HKEX filings: H1 2026 revenue will show if they can grow into 400x plus P/S
- OpenAI and Anthropic pricing: GPT-5 at $1.25 per 1M input to GPT-5.6 Sol at $5.00 shows US labs can raise prices because of distribution, Chinese labs cannot yet
- Compute tightness index at 50.3 balanced, tightening 1.9 in 30 days: watch GPU prices and token demand for Jevons paradox signals
Sources: DeepSeek API Docs pricing, OpenRouter Kimi K3 pricing, OpenAI pricing page, The Information OpenAI $25B ARR, Yahoo Finance DeepSeek theoretical $205M, Crossing River Chinese AI valuation race, @deedydas $2.6B combined, TCO model H100 costs, Compute Economics CTI.