The Kimi Panic Is Mostly a Math Error
Half the sticker price, twice the tokens burned — how Kimi K3's apparent cost advantage evaporates under scrutiny
The Kimi K3 release on July 15 sent shockwaves through Silicon Valley. A 2.8 trillion-parameter open-weights model from Chinese startup Moonshot AI, scoring 57 on the Artificial Analysis Intelligence Index and landing at #4 globally. Not just "good for a Chinese model" — genuinely competitive with GPT-5.5 and Opus 4.8.
The panic was immediate. But two sharp rebuttals cut through the noise.
On July 23, White House AI & Crypto Czar David Sacks posted: "The Kimi Panic needs to stop. American frontier models are still ahead. When you factor in what's in the lab, the gap is even larger. Let our horses run."
Four days earlier, Ben Thompson published his definitive take on Stratechery: the panic misunderstands both the economics and the technology of frontier AI.
Both make the same core argument: the cost advantage is a mirage.
The Sticker Price Trick
Kimi K3 looks dramatically cheaper on paper. But that's only half the equation.
Kimi K3 burns roughly 1.8x more tokens to arrive at the same answer. Half the sticker price × twice the tokens burned = the same bill. On the AA-Briefcase agentic benchmark, K3 averages $10.57 per task — a 10x increase from its predecessor K2.6, placing it among the most expensive models to run despite the lower per-token price.
Thompson's insight cuts deeper: this isn't a quirk of Kimi. It's the structure of the intelligence market. "A token from one model is not the same as a token from another model," he writes. "What is fungible is what is constructed from tokens, which is to say intelligence." As AI becomes a commodity, the winner is whoever has the best cost structure for serving intelligence, not whoever has the lowest token price.
The current price umbrella — where Anthropic charges $10/1M tokens blended — exists because demand for frontier intelligence still exceeds the supply of compute. When that flips, the marginal cost of serving intelligence is what matters. Thompson's bet: American labs, with months of serving experience optimizing their own best models before competitors catch up, will have that advantage.
Where Grok 4.5 Actually Stands
The 85.7% ARC-AGI-1 figure circulating for Grok 4.5 is wrong. Per ARC Prize co-creator Mike Knoop, the real numbers for Grok 4.5:
- ARC-AGI-1: 77.0% at $0.19 per task
- Performance comparable to GPT-5.4 and GPT-5.5 (Low compute)
That's strong — but it's not a record. The ARC-AGI-1 benchmark was effectively retired in December 2024 when OpenAI's o3 hit 87.5%. The current frontier is ARC-AGI-2 (where GPT-5.5 leads at 85%) and the brand-new ARC-AGI-3, where the best score is a brutal 1.86%.
Kimi K3 has no published ARC-AGI-1 score. BenchPress, a benchmark prediction tool, estimates 70.3% on ARC-AGI-2, which would place it between Opus 4.6 and GPT-5.4.
What Sacks Means by "What's in the Lab"
Sacks' claim that the gap "in the lab" is even larger isn't just rhetoric. The capabilities model shows the frontier is still accelerating: task-horizon capability has been doubling every 3.5 months since 2024, with Claude Mythos 5 currently at #1 on BenchLM (83.93). The public releases trail the internal frontier, and American labs have been shipping models at a pace no Chinese lab has matched in consistency.
The real signal, per Sacks: "Anthropic and OpenAI are growing revenue at rates that Silicon Valley has never seen before at this scale. This remains the clearest test of who is winning the market."
The Real Worry Isn't Economics
Thompson's piece closes on the actual concern. Kimi K3's weights — and now Qwen 3.8 Max's — can be downloaded and served by anyone, including bad actors. "This is a model that is approaching the frontier," Thompson writes, "and it is available for anyone to use for any purpose."
That's the cybersecurity argument. Not "China is winning AI." Not "open weights will crush American labs." The worry is that open-weights frontier models make it impossible to control who uses them and for what.
But on pure economics, the panic doesn't hold up. Intelligence markets are about the cost of serving, not the cost of downloading. And the revenue test — who's actually charging for intelligence and getting paid — still points squarely at San Francisco.
More research at bargo.ai/research.