Bargo
AI

Kimi K3 Lands: China's Moonshot Just Fired a Shot at the AI Frontier

Moonshot's 2–3 trillion parameter open-weight model launches today, landing between Claude Opus 4.8 and Fable 5. The "China is 8–12 months behind" consensus just took a direct hit.

Bargo Research

Kimi K3 is live. Moonshot AI's new flagship model began rolling out to users in mainland China late Thursday Beijing time, and is already appearing on the app for non-Chinese users. The Financial Times reported the launch was imminent only hours earlier, and the model is now available on chat.kimi.com.

K3 is a 2–3 trillion parameter open-weight model, the largest Chinese model to date. It is expected to outperform Anthropic's Claude Opus 4.8 but fall short of the newly released Fable 5. That puts it somewhere between the #1 and #2 models on the planet, depending on the benchmark, and makes it the strongest open-weight model ever released by any lab, anywhere.

The timing is not coincidental. DeepSeek's official V4 release is also scheduled for mid-July, with a new peak-time API pricing model. Two Chinese frontier labs are landing their best models in the same week. The combined message is unambiguous: the "China is 8–12 months behind" consensus is no longer sustainable.

What K3 is

Moonshot AI, a Beijing-based startup founded in 2023 by former Meta AI and Google Brain researcher Yang Zhilin, has been on a remarkable trajectory. Its first frontier model, Kimi K2, shipped only a year ago. K2 was widely praised for "taste" — a subjective measure of writing quality, creativity, and deep research capability that set it apart from the more technical, coding-focused DeepSeek models. @teortaxesTex called it "unrivaled in taste."

K2.5 added image and video support via continued pretraining on ~15 trillion tokens. K2.6 landed at #4 on the Artificial Analysis Intelligence Index with a score of 54, behind only Anthropic, Google, and OpenAI. It became the second-most used LLM on OpenRouter, behind only DeepSeek.

Now K3 appears to leapfrog directly to frontier-tier. @synthwavedd, who tested K3 early, said it "feels like another DeepSeek R1 moment" and is "often Fable level, maybe a little worse." The FT's framing: outperforms Opus 4.8, falls short of Fable 5. The full model card with benchmark tables is expected shortly.

K3 is open-weight, continuing China's deliberate strategy of open-sourcing frontier models — a strategy that has driven the open-source share of tracked inference tokens to 46%, up from 37% in late May, according to Bargo's inference economics data.

The commercial engine behind it

Moonshot is not a research lab burning venture dollars. The company:

Across the five major Chinese private AI labs, combined revenue run rate is approximately $2.6 billion, according to @deedydas. Zhipu AI (GLM) leads at $1 billion, followed by DeepSeek and Kuaishou at $500 million each, then Moonshot at $200 million. These are not science projects. They are real businesses with real revenue traction, and several are publicly traded in Hong Kong.

The DeepSeek V4 collision

DeepSeek's official V4 release is also scheduled for mid-July. The V4 preview (V4-Pro and V4-Flash) has been available since April 24, scoring roughly 52 on the Artificial Analysis Intelligence Index — below Opus 4.8 at ~61. The official release is expected to bring performance upgrades in agentic task execution, mathematical reasoning, and code generation, plus a 1-million-token context window across the lineup.

But the preview did not close the gap to the US frontier. @tphuang, a prominent DeepSeek watcher, wrote on June 30: "It's depressing for a fan of DeepSeek's work that I'm far more looking forward to the next Moonshot release than the full version of V4." @teortaxesTex was harsher after testing what appeared to be the V4 official build: "I conclude it's a failed pretrain, a failed non-scalable architecture, or DeepSeek operates on a misguided RL paradigm."

The pressure is now on DeepSeek. If V4 official ships this week and only incrementally improves on the preview, Moonshot takes the crown as China's capability leader. If V4.1 (rumored for later this summer) delivers a bigger leap, the race stays wide open. Either way, DeepSeek's new peak-time pricing — double the off-peak rate during 9am–12pm and 2pm–6pm — signals a company that is now thinking about monetization, not just scale.

What this means for the AI ecosystem

The China gap narrative is collapsing. One year ago, Chinese models were clearly behind. Today, K3 is arguably the second-best model in the world (behind Fable 5), and the open-weight leader by a wide margin. The gap has narrowed from "8–12 months behind" to "maybe a few weeks, and in some dimensions, ahead."

Open-weight is the strategic weapon. Moonshot, DeepSeek, and Alibaba's Qwen have all embraced open-weight releases. The open-source share of inference tokens has grown from 37% to 46% in less than two months. The effective blended price per million tokens has fallen to $2.39, down 21% from late May. This is the Jevons paradox in action: cheaper models are driving more demand, not less revenue.

US labs face margin pressure. Claude charges $10 per million tokens. OpenAI charges $3.50. Open-source models average $0.42. If K3 delivers Fable-adjacent performance at open-source prices, the premium pricing of US frontier models becomes harder to justify for a growing share of use cases.

GPU demand is not in crisis — yet. Bargo's Compute Tightness Index sits at 46.5 (Balanced), with H100 capacity loose, H200 tight, and B200 balanced. The index has been flat for two months. A flood of new K3 inference demand could change that, but for now, the GPU rental market is absorbing the open-source boom without strain.

What to watch


More research at bargo.ai/research.

Get Bargo research in your inbox
One email when we publish. No spam, unsubscribe anytime.
Get Bargo research in your inbox
One email when we publish. No spam.