Moonshot AI just put out the biggest Chinese open-source model ever released, and it topped Claude Fable 5 at writing scripts.
K3 also claimed the top spot on Arena AI's Frontend Code Leaderboard—a ranking built from thousands of pairwise human votes on code generation tasks, again Elo-scored—with 1,679 against Fable 5's 1,631. First place in six out of seven frontend domains.
This is a 17-place jump from Kimi-k2.6 (#18 -> #1).
Kimi K3 is beating Fable 5 head to head on BridgeBench.
Same task, blind judge panel.
K3 has won 7 of 8 arenas, including Refactoring 9-0 and Debugging 6-1.
Fable 5's only win: Speed.
This is the first time that an open model is ahead of all proprietary ones for this comprehensive web engineering benchmark.
Notes:
What this thing actually isK3 packs 2.8 trillion parameters—the numerical values that store a model's knowledge—in a mixture-of-experts architecture. Mixture of experts splits those parameters into 896 "expert" subnetworks and activates only a fraction for any given task. That's how you get frontier-level intelligence without melting the server room.

It comes with a one-million-token context window—tokens are the basic unit of information an AI processes, about three-quarters of a word each—native image and video understanding, and always-on reasoning.
Two architectural techniques underpin the efficiency gains. Kimi Delta Attention speeds up decoding for long sequences—up to 6.3x faster at million-token contexts. Attention Residuals routes information selectively across model layers rather than accumulating it uniformly, adding about 25% training efficiency at under 2% extra compute cost—together yielding roughly 2.5x better scaling efficiency than K2.
Benchmarks are nice, Prices are nicerKimi K3 costs $3 per million input tokens and $15 per million output tokens—the same rate as Claude Sonnet 5, Anthropic's mid-tier model. The difference is that Sonnet 5 is Anthropic's middle-ground offering; K3 is sitting three points below Fable 5 on the Artificial Analysis composite. Per task across that nine-benchmark suite, K3 runs $0.94 versus $1.04 for GPT-5.6 Sol and $1.80 for Opus 4.8.
In other words, this model offers top of the line performance at mid-tier level prices.
If Anthropic goes ahead with its intentions of making Fable 5 available only via API, K3 becomes the nearest open-weight alternative to whatever model currently sits second in the industry—at half the per-task cost of Opus 4.8. That's the scenario benchmark chasers are already running the math on.
K3's launch is the argument U.S. chip export controls advocates don't want to have. The U.S. restricted Nvidia's H800 GPUs from export to China in late 2023; Moonshot confirmed it trained earlier models on those chips. K3's own benchmark documentation references H200s and what the company calls "a GPGPU from an alternative vendor"—widely interpreted as Huawei Ascend hardware—without specifying where that hardware sits.
The asterisk you should readK3's hallucination rate on AA-Omniscience—a benchmark that measures how often a model confidently fabricates an answer it doesn't know—jumped from 39% to 51% compared to predecessor K2.6. More correct answers overall; more made-up ones too. The model also acknowledges in its own documentation that it can be "excessively proactive," making unexpected decisions on a user's behalf during long autonomous tasks.
If you want to try it for free, you can. It’s available on Kimi’s official website. But good luck: The servers are so packed that tasks get interrupted constantly due to traffic constraints, making it barely usable. A better alternative is to either pay for a subscription or use it over an API.

Weights will be released on July 27. Those will be available for big enterprises and businesses. No domestic GPU, no matter how big, is currently able to handle a model this size.


















