Kimi K3 评测:最强开源模型,但落后闭源前沿数月
On Kimi K3: Its Capabilities And Related Discontents
Kimi K3 is a very good model with excellent benchmarks. Assuming its weights are released as planned it will become, purely in terms of raw capability, the strongest open model.
Do not get carried away. Do not judge Kimi K3 only its relative strengths. In aggregate it is several months behind the closed model frontier, at least four and my median guess is six, with the post-training closer and the pre-training farther out. This is less months than before, but the months are denser now.
It is somewhat distilled. It likely outperforms on benchmarks relative to practical performance. All its benchmarks are scored at maximum effort, typically a lot more tokens than are used in similar tests by Fable or Sol. Performance looks jagged. Kimi will be excellent at some things, less so at other things.
We will know more over the coming weeks. For now access is spotty and not that many people have actually had the chance to try Kimi K3, so I have larger error bars than usual around its capabilities. Alas, time waits for no one, so we press on.
It is the largest open model so far at 2.8T, on the upper end of possible sizes for Claude Opus and near the bottom of possible sizes for Mythos, which explains many of its gains. It is slow and appears hungry for tokens. A lot of the positive reactions are to this being a big model, which thus has at least a decent amount of ‘big model smell’ and generally trades being slower and more expensive for some performance gains. That’s a great move, but it should be a while before they can do it again.
Distillation from Claude is clearly part of the story, likely largely from Fable, and is clearly nothing close to the whole story. Clearly Moonshot would do this even if it helped only a little. The timing of them releasing a bigger model is suggestive.
It is again a good model, but once you correct for the overperformance on benchmarks and look at expected practical performance, it is not clear this is so different from what you would have expected from a Kimi K3 that had 2.8T parameters.
Consider that Kimi K3’s (preliminary unofficial) Epoch Capabilities Index is exactly on the Chinese trend line.
Kimi K3 is absolutely worth checking to see if it fits into your workflows. At this price point, for both the API and the subscription, it is not going to fill the role of the smaller cheaper open models, and I expect it to usually lose out in a fight with the top closed models, but there are going to be some places where Kimi K3 is a good choice.
Andrew Curran: Following the success of Kimi K3, Moonshot AI has informed investors that it plans an IPO in Hong Kong within the next six months, according to Bloomberg.
Good idea. Strike while the iron is hot.
Table of Contents
- DeepSeek Moments: Here We Go Again.
- We Had a Moment (Reprise from June 2025).
- The Story Since Then.
- The Kimi K3 Announcement, Pitch and Basic Facts.
- On Modern Benchmaxxing.
- Other People’s Benchmarks.
- Benchmarks Are Not The Real World.
- Technical Safeguards? What Are Those?
- Things Kimi Can Do.
- Things Kimi Cannot Do.
- Things It Is Not Easy To Get Kimi To Do.
- Open Weight Models Are Unsafe And Nothing Can Fix This.
- Dean Ball Attempts To Be Constructive.
- OpenAI Employees Are Relatively Bullish On This One.
- Kimi K3 Is Relatively Strongest At Typical Agentic Coding and 3D.
- Reactions.
- Who Are You?
- How Did They Do It?
- Conclusion.
DeepSeek Moments: Here We Go Again
All discourse about Chinese models lives in the shadow of the DeepSeek moment.
There are a lot of people who really, really want another DeepSeek moment to happen.
These people really, really want to tell the story that Chinese open models are catching up to American closed models, that AI and inference will become commoditized.
Their motivations vary. They often want to affirm open models, or the importance of the ‘tech stack.’ Others simply want to see OpenAI and Anthropic go down, or know that such claims sell. Often the ultimate objective is to argue against all AI regulations, or anything that might ‘slow us down’ or cause us to ‘lose to China.’
It is actively suicidal to respond to ‘the Chinese have better models now’ with ‘then we had better sell them the compute so they can run them and also build even better ones.’ Yet every time, yes, people will argue that. Sigh.
Often they simply want to tell American AI to stop taking precautions, to stop being annoying and take down the classifiers, as in ‘genie is out of the bottle, so release the bigger genie with unlimited wishes, it’s the only way.’ People really would take major catastrophic risks rather than deal with classifiers, and are Big Mad about this.
Google was down 4.4% on the day, SpaceX was down 3.1% and Nvidia down over 2%, and tech stocks were down again on Friday, so plausibly we’re doing this again.
And yep, we are at risk of doing this again:
Axios (being wrong): China just erased America’s AI lead
Seán Ó hÉigeartaigh: No it didn’t. (although Kimi K3 is undoubtedly impressive)
The pattern is:
- Chinese model releases.
- There is some impressive benchmark cited.
- Therefore, America’s lead is gone, QED, that’s it, no, really, that’s it.
In the Axios case the benchmark in question is Arena. Based on that alone, they state as fact that America’s lead is gone.
I would ignore, but this style of logic has convinced a lot of Washington D.C. multiple times, and that has had substantial policy impact.
I shudder to think what such folks might do now that they also know about Mythos. The confusion over Fable jailbreaks could easily extend to a broader dumb panic.
The original DeepSeek moment happened because of a confluence of events.
Let’s review.
We Had a Moment (Reprise from June 2025)
We all remember The DeepSeek Moment, which led to Panic at the App Store, lots of stock market turmoil that made remarkably little fundamental sense and that has been borne out as rather silly, a very intense week and a conclusion to not panic after all.
Over several months, a clear picture emerged of (most of) what happened: A confluence of narrative factors transformed DeepSeek’s r1 from an impressive but not terribly surprising model worth updating on into a shot heard round the world, despite the lack of direct ‘fanfare.’
In particular, these all worked together to cause this effect:
- The ‘six million dollar model’ narrative. People equated v3’s marginal compute costs with the overall budget of American labs like OpenAI and Anthropic. This is like saying DeepSeek spent a lot less on apples than OpenAI spent on food. When making an apples-to-apples comparison, DeepSeek spent less, but the difference was far less stark.
- DeepSeek simultaneously released an app that was free with a remarkably clean design and visible chain-of-thought (CoT). DeepSeek was fast following, so they had no reason to hide the CoT. Comparisons only compared DeepSeek’s top use cases to the same use cases elsewhere, ignoring the features and use cases DeepSeek lacked or did poorly on. So if you wanted to do first-day free querying, you got what was at the time a unique and viral experience. This forced other labs to also show CoT and accelerate release of various models and features.
- It takes a while to know how good a model really is, and the different style and visible CoT and excitement made people think r1 was better than it was.
- The timing was impeccable. DeepSeek got in right before a series of other model releases. Within two weeks it was very clear that American labs remained ahead. This was the peak of a DeepSeek cycle and the low point in others cycles.
- The timing was also impeccable in terms of the technology. This was very early days of RL scaling, such that the training process could still be done cheaply. DeepSeek did a great job extracting the most from its chips, but they are likely going to have increasing trouble with its compute disadvantage going forwards.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力