China Kimi K3 is out – and beats Claude Fable and GPT 5.6 Sol in key benchmarks


short

  • Moonshot AI launched the Kimi K3 on July 16 – a 2.8 trillion parameter open-weighted model that outperforms US labs on specific specialized benchmarks.
  • K3 is priced similarly to Claude Sonnet 5 ($3 per million input codes, $15 per million output codes) while scoring closer to Fable 5.
  • Full model weights—the files that allow anyone to run, improve, or build on the model locally—drop by July 27 under a modified license from MIT, making K3 the largest freely available AI model in history.

Moonshot AI has just released the largest open source Chinese model ever released, and it has even surpassed Claude’s Fable 5 in scripting.

Towards elo writing for artificial intelligence– A benchmark where models write real scripts that are blindly judged against published versions, and are scored using the same Elo system that ranks chess players – Kimi placed K3 at 2840, higher than Fable 5 (maximum) at 2760. This is a rating that the Anthropy team has historically dominated.

K3 also took first place on Arena AI’s Frontend Code Leaderboard — a ranking created from thousands of human pairwise votes on code generation tasks, again an Elo record — with a score of 1,679 versus Fable 5’s 1,631. First place in six out of seven frontend areas.

the Artificial Analysis Intelligence Index– a score based on nine independent ratings covering programming, reasoning, agentic action and knowledge, with a score of 0 to 100 – puts the K3 at 57, with the Claude Fable 5 at 60, the GPT-5.6 Sol at 59, and the Claude Opus 4.8 at 56. This places the K3 as the third most capable model on the composite, with the Fable 5 only beating it by 3%.

If you want to get an idea of ​​what it can do, this is it Zero result A prompt that asks the form to create an iOS version. For comparison, this It is the best approximation shared on social media using GPT 5.6 Sol and a more detailed claim.

What actually is this thing?

K3 combines 2.8 trillion parameters — numerical values ​​that store model knowledge — in a mixture of experts. Expert Mix divides these parameters into 896 “expert” subnetworks and activates only a small portion for any given task. This is how you get border-level intelligence without melting the server room.

“It is the world’s first open source model in the 3 trillion parameter class, designed for frontier intelligence scenarios including long-term programming, cognitive work, and reasoning,” Moonshot AI He says. This isn’t marketing theatre: DeepSeek’s V4-Pro tops out at 1.6 trillion parameters, while Moonshot’s K2 tops out at a trillion. The K3 nearly doubles the closest open weight competitor on the size chart.

It comes with a contextual window containing a million tokens – tokens are the basic unit of information processed by AI, about three-quarters of a word each – understanding original images and video, and always thinking.

There are two architectural approaches that support efficiency gains. Kimi Delta Attention accelerates decoding of long sequences, up to 6.3x faster in 1 million token contexts. Attention Residuals selectively routes information across model layers rather than collecting it uniformly, adding approximately 25% training efficiency at less than 2% additional compute cost – together achieving approximately 2.5x better scaling efficiency than K2.

The standards are nice, and the prices are nicer

The Kimi K3 costs $3 per million input codes and $15 per million output codes — the same price as the Claude Sonnet 5, Anthropic’s mid-level model. The difference is that Sonnet 5 is a medium-anthropic offering; K3 is three spots below Myth 5 on the Synthetic Analysis Composite. Per task across this set of nine benchmarks, K3 is priced at $0.94 versus $1.04 for GPT-5.6 Sol and $1.80 for Opus 4.8.

In other words, this model offers the highest level of performance at mid-level prices.

like Decryption covered in Maythe pricing gap between Chinese and US frontier AI was between 15 and 30 times earlier this year. The K3 is no less expensive than DeepSeek—it’s more like a mid-range Western model—but it offers near-bounds performance at this level. For API-driven teams, this represents a significant cost improvement.

If Anthropic goes ahead with its intentions of making Fable 5 available only via the API, the K3 becomes the closest open-weight alternative to any model currently ranked second in the industry – at half the cost per mission of the Opus 4.8. This is the scenario that standard chasers are already doing the math on.

The K3 launch is the argument advocates of US chip export controls don’t want to have. US restricts export of Nvidia H800 GPUs to China in late 2023; Moonshot confirmed that it had trained previous models on these chips. The K3’s benchmark documentation refers to the H200s and what the company calls a “GPGPU from an alternative vendor” — which is widely interpreted as a Huawei Ascend device — without specifying where that device is located.

Moonshot AI’s president, Yutong Zhang, laid out the framework for the restrictions live in Davos this year Silicon Republic: “We knew we didn’t have the luxury to simply scale computing… This forced us to focus on basic research and efficiency.” Bank of America analysts wrote, in a note after the launch, that K3 proves that “expansion of pre-training, coupled with architectural innovation, can still deliver step-change gains for leading Chinese models” under these constraints.

Moonshot is one of the so-called AI Tiger startups That collectively changed the global modeling landscape without access to the chips Washington said it would need. Whether this is an argument for tightening export controls or an argument that they are ineffective is a political question that Washington has not decided.

The star you should read

AA-Omniscience’s K3 hallucination rate — a metric that measures how often a model confidently makes up an answer it doesn’t know — jumped from 39% to 51% compared to the previous K2.6 model. More correct answers overall; More of them are made up too. The model also acknowledges in its own documentation that it can be “overly proactive”, making unexpected decisions on behalf of the user during long autonomous tasks.

For the teams that managed Tools based on Kimi K2.6 And want to upgrade, the K3 is a meaningful step up on most fronts – but this hallucinatory Delta is worth a stress test before you trust it with anything needing precision.

If you want to try it for free, you can. It is available on the official Kimi website. But good luck: the servers are so congested that tasks are constantly interrupted by traffic restrictions, making them barely usable. A better alternative is to either pay for the subscription or use it via the API.

Weights will be released on July 27. These will be available to larger institutions and companies. No domestic GPU, regardless of size, is currently capable of handling a model of this size.

Daily debriefing Newsletter

Start each day with the latest news, plus original features, podcasts, videos and more.





Source link

Leave a Reply

Your email address will not be published. Required fields are marked *