DeepSeek's Leaked Playbook: Give Away the Model, Own the Inference
The Liang Wenfeng transcript reveals a 10-month GPU payback at 85% margins — and open-source just captured three-quarters of all AI token volume
The transcript was never supposed to go public. In May, DeepSeek founder Liang Wenfeng sat down with a small group of potential investors for what became a nearly four-hour meeting. The recording leaked last week. DeepSeek did not officially confirm the document's authenticity, but the company put its second funding round on "indefinite hold" immediately after — which, in the world of Chinese AI, is about as close to a confirmation as you get.
Convequity called it immediately: this transcript reveals "the real business model for open-source AI." The framework is simple and devastating: give away the model weights, monetize the inference. Open-source distribution becomes a marketing funnel. The proprietary inference stack — which only the original lab fully understands — captures the margin.
The transcript makes this explicit. Liang says DeepSeek recovers hardware costs in roughly ten months, despite depreciating servers over three to five years. Inference margins run around 85%. The company could raise prices by 50% to 100% without materially reducing demand. They choose not to. "This is not profit-maximizing pricing," he says. "Our purpose is to make this useful to people, rather than to make the most money."
The Numbers: What 35x Cheaper Actually Looks Like
The price gap between DeepSeek and the US frontier labs has stopped being a curiosity and become the story. Here is the live pricing as of today:
| Provider | Model | Input $/1M tokens | Output $/1M tokens |
|---|---|---|---|
| DeepSeek | V4 Flash | $0.14 | $0.28 |
| DeepSeek | V4 Pro | $0.44 | $0.87 |
| OpenAI | GPT-5.6 Sol | $5.00 | $30.00 |
| OpenAI | GPT-5.6 Terra | $2.50 | $15.00 |
| Anthropic | Claude Opus 5 | $5.00 | $25.00 |
| Anthropic | Claude Sonnet 5 | $2.00 | $10.00 |
Sources: DeepSeek, OpenAI, Anthropic. Sonnet 5 pricing is introductory through August 31, 2026, after which it rises to $3/$15.
V4 Flash is roughly 35 times cheaper on input and 107 times cheaper on output than GPT-5.6 Sol. Even DeepSeek's premium V4 Pro tier is about 11 times cheaper on input and 29 times cheaper on output than the US frontier flagships. With cache hits — common in production workloads — the gap widens to beyond 100x.
The Jevons Effect Is Already Here
Cheaper inference does not shrink the market. It expands it. The Bargo Token Demand Index tracks this in real time. Open-source providers, led by DeepSeek, now claim 74.6% of all tracked token volume, up from roughly 60% two months ago. Token consumption grew 20% month over month, while the effective blended price fell to $1.62 per million tokens. The total market — all providers combined — consumes about 55 trillion tokens per week, and the tokenomics model projects this growing at roughly 2.5 times annually.
This is the Jevons paradox in AI: demand volume is rising faster than prices are falling. Total implied spend — list prices times token volume — is not shrinking. It is growing. The GPU rental market confirms the supply side is feeling the pressure: the Compute Tightness Index sits at 54 (Balanced), but has tightened 10 points in the last 30 days, and H200 capacity is already in the Tight regime.
The Inference Moat: Why Open Weights Don't Equal Commodity Inference
The most important claim in the transcript is not about pricing. It is about cost structure. Liang says DeepSeek achieves 85% inference margins because its optimization is substantially better than competitors. "Without our optimizations, the costs for companies such as Alibaba or Tencent would probably be several times higher."
This is the thesis in one sentence. The model weights may be open, but the inference stack is a black box. The lab that built the model understands its architecture, kernel behavior, memory access patterns, and quality-performance trade-offs better than any third-party provider ever will. An open-weight model served by Together or Fireworks cannot match the economics of the same model served by DeepSeek on DeepSeek's own optimized stack.
The Bargo Cloud TCO model provides a sanity check. At full cost, serving an H100 cluster costs about $0.074 per million tokens. At cash cost — excluding capex recovery and financing — it drops to $0.025. DeepSeek's V4 Flash pricing at $0.14 input and $0.28 output is comfortably above both figures, even before factoring in the inference optimizations Liang claims. The 10-month hardware payback math is aggressive but directionally credible.
Liang puts the breakeven point at $1 billion in annual API revenue. The Information recently reported DeepSeek's ARR approaching $500 million. The company had roughly 20,000 Hopper-equivalent GPUs as of May, and Liang says the plan is to spend "basically all NVIDIA" over the next six months — constrained only by what export controls allow.
The Strategic Logic: Land-Grab, Not Generosity
The instinct is to read DeepSeek's pricing as altruism. It is not. This is a classic platform land-grab.
At 85% margins, DeepSeek can price at levels where it earns a healthy return while competitors lose money on every token served. Open weights ensure the model is everywhere. The proprietary inference stack ensures DeepSeek captures the economics. Liang could raise prices — he says as much — but raising prices would invite competition and slow adoption. Low prices accelerate the flywheel: more users, more tokens, more optimization data, better models, even more users.
The Bargo tokenomics model shows the global inference revenue pool reaching roughly $380 billion annualized by Q4 2026 and over $1 trillion by Q4 2027 in the base case. DeepSeek capturing even 15% to 20% of that pool at 85% margins is a business that justifies essentially any GPU budget.
The Hardware Reality Check
The transcript is clear-eyed about DeepSeek's constraints. Liang admits the company has secured two Huawei Atlas Super Clusters — 16,000 Ascend 950DT accelerators — and is porting TileLang to Huawei's software stack to break the CUDA dependency. But he is blunt about the alternative: "If B200 were available, any quantity at reasonable prices would be worthwhile." The constraint is availability, not willingness to pay.
He also notes that Huawei chips show meaningful efficiency degradation by the third year of use. The 10-month payback math works on NVIDIA hardware. On domestic Chinese silicon, the unit economics are materially worse. This is the real risk to the strategy: export controls forcing DeepSeek onto inferior hardware eventually eat the inference margin advantage.
As we detailed in the DeepSeek chip gambit analysis, the hardware reality means DeepSeek's open-source playbook is powerful but not unconstrained. SMIC's 7nm process cannot close the gap. No HBM access. Liang's own admission that "if we could buy NVIDIA cards, domestic substitution would be hard" sets a hard ceiling on how far the inference moat can extend on Chinese silicon.
What This Means
The leaked transcript validates what Gavin Baker and Chamath Palihapitiya have been arguing from different angles: open-source AI is not an ideology. It is a business strategy. And it is working.
For US frontier labs, the strategic question has sharpened. OpenAI and Anthropic charge 35x to 100x more for models that are better — Claude Opus 5 and GPT-5.6 Sol still lead on the hardest reasoning benchmarks — but the gap is narrowing. When DeepSeek's V4 Flash can serve most production workloads at less than one percent of the US frontier price, the premium for "best model" is under constant compression.
For NVIDIA, the transcript is a pure demand signal. Liang wants to buy every B200 he can. The constraint is policy, not willingness. A Chinese lab with 75% of global token volume and a stated goal of spending "every cent" on GPUs is the most bullish possible read on AI hardware demand — as long as the US does not close the door entirely.
For investors, the framework is now explicit. The AI pricing fork that opened in early 2026 has widened. Open-source inference at DeepSeek's economics is a structural force, not a cyclical one. The companies that benefit are the compute platform owners. The companies that face pressure are the premium inference providers charging 35x the market-clearing price.