跳到主内容
@wquguru
精选88Import AI(RSS)行业动态多源精选 ×10

Kimi K3 缩小中美模型差距,开源模型网络能力逼近闭源

Import AI 465: Open vs closed gaps; Kimi K3; Demis’ big policy plan

原文
发到 X
推荐理由

三则重磅动态:AISI 开源-闭源差距报告、Kimi K3 逼近前沿、Demis 提出 AGI 监管框架。做安全、政策或模型研究的同学建议细读原文。

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe.

Subscribe now

UK government: Gap between open and closed weight models on cyber is shrinking:

…The cyber-eschaton cometh…

The UK government’s AI Security Institute (AISI) has analyzed the delta in cybersecurity capabilities between powerful proprietary models and open weight models. The results show that this year, the gap has shrunk. “This is our first public analysis of how far leading open weight models trail the closed cyber frontier,” AISI writes. “Recent open models GLM-5.2 and DeepSeek V4-Pro perform similarly to frontier closed models released 4 to 7 months before them – a narrower gap than the 6 to 10 months we measured through most of 2025.”

Specific details: On a set of 70 evals for specific, narrow cyber capabilities, GLM-5.2 is closest to Claude Opus 4.6, which was released 4.3 months earlier, while DeepSeek-V4-Pro sits somewhere between Claude Opus 4.5 and GPT-5 (released in November and August 2025, respectively). “AISI intends to test Kimi K3 on this same basis, once its weights are publicly released,” AISI writes.

The gap lengthens a bit for long-horizon cyber ranges, which are tasks that see how well models can chain various capabilities together to complete a full hacking operation. Specifically, on a cyberrange called The Last Ones, “GLM-5.2 reaches as far as Opus 4.5, a model released less than 7 months before it, while DeepSeek’s V4-Pro falls below Sonnet 4.5 (a sub-cyber-frontier model released 7 months before it),” AISI writes. “The gap here is larger than on our narrow cyber tasks”.

This, I think, rhymes with the idea that though open weight models can be superficially quite strong, they sometimes lack a bit of the generalization magic juice that distinguishes proprietary models. This is what people in the AI industry call “big model smell”.

Why this matters – the offense and defense balance of the world is about to change: The main implication here is that the gap between the controllable frontier and the lawless openly diffused frontier is shrinking. “This implies cyber defenders have a short window to prepare before today’s frontier cyber capabilities may become accessible without the same safeguards” used by proprietary companies, AISI writes.

Read more: How Far Behind the Frontier are Leading Open Weight Models on Cyber? (UK AI Security Institute blog).

Kimi: China shortens the gap between Chinese and Western models:

…Plus, early signs of AI R&D…

In the last couple of years, Chinese firms have begun to out-compete Western actors at building and deploying open weight models (e.g, DeepSeek), and now are starting to close the gap on frontier models as well. The latest and best example of this is Kimi K3, a 2.8 trillion parameter model. Kimi has exceptionally strong scores on all the tasks that the major proprietary ones benchmark on and typically matches or trails Claude Fable 5 and GPT 5.6 Sol.

However, Kimi has some brittleness which smells to me like “benchmaxxing” – performance may have been tuned around these benchmarks in a way that harms some parts of generalization.

“While its overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol, Kimi K3 demonstrated frontier-level performance across our evaluation suite, consistently outperforming other tested models,” Kimi writes.

Kimi’s weights will be made available in the coming weeks along with a research paper about the model.

AI that builds AI: Kimi has some example use-cases which relate to recursive self-improvement; using AI systems to improve AI itself. Specifically, they tested out how good Kimi was at writing GPU compilers. “Kimi K3 developed MiniTriton, a compact Triton-like compiler with its own tile-level IR layer over MLIR, optimization passes, and a PTX code-generation pipeline. Across supported roofline benchmarks, MiniTriton delivers performance on par with or better than Triton and torch.compile — beating Triton on certain workloads,” they write. (Though note they don’t talk about any of this stuff going into actual production, aka being used to train Kimi K3 itself, but it’s certainly suggestive that future models might be able to do this.)

Additionally, they showed how “Kimi K3 designed a chip to serve a nano model built on its own architecture. In a single 48-hour autonomous run, K3 built, optimized, and verified the chip using open-source EDA tools on the Nangate 45nm library.”

Why this matters – widely diffused AI systems are getting a lot better: Most notions of AI policy and AI safety rest on control – the idea that there’s a small number of actors deploying proprietary models which you can intervene on at the platform level (e.g, via classifiers or know your customer gates) alongside the model level. Models like Kimi K3 – if they go through with releasing the weights – completely change this by diffusing broadly uncontrollable powerful AI into the world. This will have a vast range of positive effects, driving a boom in entrepreneurship and increasing the ‘sovereign intelligence’ available to anyone who can run the model, but will also yield various unknown unknowns. The next few years are going to be defined by the gap between proprietary models and widely available models and how they show up in society will determine much of the policy discussion.

Read more: Kimi K3: Open Frontier Intelligence (Kimi blog).

Demis Hassabis proposes a regulatory regime for artificial general intelligence:

…FINRA for AI…

DeepMind founder Demis Hassabis has laid out a policy prescription for AGI. His basic idea is that the US government should develop a framework for testing out frontier AI systems for new capabilities and should do this via a Standards Body modelled on a federally overseen public-private partnership or self-regulatory organization, much like the Financial Industry Regulatory Authority (FINRA). “This US-initiated effort would provide a strong starting point for creating shared international standards on Frontier AI,” he says.

What the standards body would do: “The Standards Body would be responsible for developing assessment protocols and working with appropriate federal agencies and the US National Labs to conduct testing in areas relevant to national security,” Hassabis writes. This testing infrastructure would help to define what would make a model a “Frontier Model”, and labs developing those models would “be encouraged” to adopt best practices in areas like publishing details about their systems, investing in cybersecurity, personnel vetting, and more.

Start voluntary and then move to law: “Initially, Frontier Labs would voluntarily share models with the Standards Body for review up to 30 days before release. Once the assessment protocol is shown to be effective and robust, formalisation could quickly follow,” he writes.

Why this matters – emerging industry consensus: Demis’ piece is interesting because it pulls together some de facto consensus positions that have emerged across the AI industry in recent years; powerful AI systems should be tested by third parties that have some loose relationship to a regulator (e.g, the US government). It also rhymes with the de facto policy norm that has emerged in America recently across both the Trump admin’s executive order about AI as well as the recent processes developed in the aftermath of the Anthropic export controls saga; here, government and industry developed assessment methods for evaluating the capabilities of AI systems and figuring out if they posed national security risks.

Demis’s piece is also interesting because Google is rarely this forthright about policy – this is a reassuringly specific proposal and it sits alongside spiritually similar proposals from Anthropic (albeit somewhat toothier).

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近