Bargo
AI

Kimi K3 Is a Compute Hog. That's Bullish.

The per-token price debate misses the point. K3 is a 2.8T-parameter brute with an immature inference stack that burns three times the tokens of its peers. The open-weight release today is a compute-demand story.

Bargo Research

Kimi K3 landed two weeks ago at $3 per million input tokens and $15 per million output — 70% below Claude Fable 5's $10/$50. The narrative wrote itself: China has a frontier model, and it is cheaper. A second wave of analysis pushed back, pointing out that K3 burns through so many tokens per task that it actually costs more on certain benchmarks.

Both sides are arguing about the wrong thing. The price debate is a distraction. The real story is that K3 is a 2.8-trillion-parameter model with an immature inference stack that consumes three times the tokens of its peers on agentic workloads. Its open-weight release today means more developers running it, burning more tokens, on more GPUs. The pricing question is a footnote. The compute demand is the thesis.

The numbers behind the price debate

K3's cost per task varies by 10x depending on what you ask it to do. The same model is the most expensive frontier model on one benchmark and the cheapest on another.

Benchmark What it tests Kimi K3 GPT-5.6 Sol Claude Fable 5
AA-Briefcase Agentic knowledge work $10.57 (120k tokens, 83 turns) $1.04 (42k tokens, 50 turns) $3.25 (42k tokens, 67 turns)
DeepSWE Software engineering $4.65 (68.5% pass@1) $8.39 (73% pass@1) $13.41 (69.9% pass@1)
Composio SaaS agent workflows $1.39 (463k tokens, 7/12) $2.69 (538k tokens, 6/12) $7.76 (776k tokens, 7/12)

On Artificial Analysis' AA-Briefcase, K3 is the most expensive frontier model. It averages 83 turns and 120,000 output tokens per task — nearly three times the tokens of Fable 5 or GPT-5.6 Sol — and takes 56 minutes per task. It costs $10.57 per task versus $3.25 for Fable 5 and $1.04 for GPT-5.6 Sol.

Cost per AA-Briefcase Task (agentic knowledge work)

On Together AI's DeepSWE analysis, the picture reverses. Across 452 graded software engineering rollouts, K3 costs $4.65 per task at 68.5% pass@1, versus Fable 5 at $13.41 and 69.9%. Per solved task, K3 delivers 2.8x more work per dollar. It also pulls ahead on pass@2 (82.0% vs 80.2%) and pass@4 (89.4% vs 88.5%).

Cost per DeepSWE Task (coding)

On Composio's 12-case SaaS agent suite — real account-backed CRM, calendar, and GitHub workflows — K3 used 463,000 tokens per case versus Fable's 776,000. It tied Fable's 7/12 score for a sixth of the price.

The price story is not that K3 is cheap or expensive. It is that the variance is enormous, and the variance is a product of an immature inference stack, not a pricing strategy.

Why K3 burns so many tokens

K3 cannot turn off its reasoning. It has low, high, and max effort modes, but even "low" still does chain-of-thought. Fable 5 and GPT-5.6 Sol can toggle thinking off entirely for simple steps. On a multi-turn agent run, that reasoning tax compounds. On AA-Briefcase, K3 takes 83 turns per task — the other models take 50 to 67 — and each turn burns tokens that its peers skip.

SemiAnalysis noted that the open-weight release was delayed to today, July 27, partly to let vLLM and SGLang optimize serving so the model would not look broken at launch. The inference stack is raw. The token burn is not a feature. It is the visible edge of a three-year-old company's first attempt at serving a 2.8-trillion-parameter model.

The compute demand is the story

The DeepSeek analogy — a model that does more with less compute — is backwards. K3 is a scaling play. It proves that the brute-force approach still works, and it does so at a parameter count that demands more hardware, not less. SemiAnalysis noted it is too large to fit on a single DGX B200. The open-weight release means the model will be fine-tuned, forked, and deployed across the open-source ecosystem. Every deployment burns tokens. Every fine-tune burns training compute. The pricing debate is a sideshow. The real story is that the frontier is still hungry, and the open-weight release feeds it.

Get Bargo research in your inbox
One email when we publish. No spam, unsubscribe anytime.
Get Bargo research in your inbox
One email when we publish. No spam.