Mistral发布1.05T参数多模态MoE模型ML4
Mistral AI Releases Mistral Large 4 (Le Chonk): A 1.05T Parameter Multimodal MoE Model
Mistral旗舰大模型更新,1.05T参数规模配合Grace Blackwell硬件训练,性价比突出,值得关注其安全能力突破。
Mistral AI has just announced the release of Mistral Large 4 (ML4), internally nicknamed Le Chonk, as a public preview. ML4 is a granular Mixture of Experts model with 1.05 trillion total parameters, 49 billion active per token, a 1.6 billion parameter vision encoder, and a 1 million token context window, per the model documentation. It was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own European datacenters.
Mistral AI 刚刚宣布发布 Mistral Large 4(ML4),内部代号为 Le Chonk,作为公开预览版。根据模型文档,ML4 是一个细粒度的混合专家(MoE)模型,总参数量为 1.05 万亿,每 token 激活 490 亿参数,配备一个 16 亿参数的视觉编码器,以及 100 万 token 的上下文窗口。它是在 Mistral 位于欧洲的自有数据中心中,使用 3,800 块 NVIDIA Grace Blackwell GPU 从头训练而成的。
TL;DR
TL;DR(太长不看版)
Mistral AI released Mistral Large 4 (‘Le Chonk’) as a public preview on 6 October 2026: a 1.05T parameter granular MoE with 49B active per token, native image input, and a 1M context window, trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own EU datacenters. The API is live now at $1.36 per 1M input and $4.18 per 1M output tokens, but the weights do not ship until end of October, so self hosting is not yet possible. Its standout results are in cybersecurity, where Mistral reports 93% on Cybench and 82% on CyberGym-E2E and notes that several closed frontier models score near zero because they refuse the task.
Mistral AI 于 2026 年 10 月 6 日以公开预览形式发布了 Mistral Large 4(‘Le Chonk’):这是一个拥有 1.05T 参数的细粒度 MoE 模型,每 token 激活 49B 参数,支持原生图像输入,上下文窗口为 1M,在 Mistral 自有的欧洲数据中心中使用 3,800 块 NVIDIA Grace Blackwell GPU 从头训练而成。API 现已上线,输入 token 价格为每 1M 美元 1.36,输出 token 价格为每 1M 美元 4.18,但权重文件直到 10 月底才会发布,因此目前尚不支持自托管。其突出表现集中在网络安全领域,Mistral 报告其在 Cybench 上得分 93%,在 CyberGym-E2E 上得分 82%,并指出由于拒绝执行任务,几个封闭的前沿模型在该测试中得分接近零。
What is the architecture?
架构是什么样的?
ML4 is a hybrid instruct-and-reasoning MoE that takes image input natively. Only about 4.7% of the weights activate per token, which is how a 1 trillion class model serves at mid-tier pricing. The full 1.05T still has to sit in memory, so the activation count sets compute, not your hardware bill.
ML4 是一种混合指令与推理能力的 MoE 模型,原生支持图像输入。每个 token 仅激活约 4.7% 的权重,这使得一个拥有万亿级参数的模型能够以中等层级的价格提供服务。完整的 1.05T 参数仍需驻留在内存中,因此激活数量决定了计算量,而非你的硬件账单。
Mistral has not yet published the expert count, top-k routing, or layer layout; those arrive with the weights. Training data spanned more than 160 languages, including every official EU language.
Mistral 尚未公布专家数量、top-k 路由策略或层级布局;这些细节将在权重发布时一同提供。训练数据涵盖了 160 多种语言,包括所有欧盟官方语言。
Interactive Explainer
交互式解释器
How does it perform?
性能如何?
In cybersecurity, Mistral reports 93% on Cybench and 82% on CyberGym-E2E, placing ML4 in the global top 5 on the Artificial Analysis Cyber Index. The more interesting claim is structural: Mistral states several frontier closed models score near zero on CyberGym-E2E because they refuse outright. Reproducing a vulnerability to prove it is real is standard defensive work, and provider-level refusals block it.
在网络安全方面,Mistral 报告其在 Cybench 上得分为 93%,在 CyberGym-E2E 上得分为 82%,使 ML4 在 Artificial Analysis 网络安全指数全球排名前五。更有趣的说法是结构性的:Mistral 指出,几个前沿封闭模型在 CyberGym-E2E 上得分接近零,因为它们直接拒绝执行任务。复现漏洞以证明其真实性是标准的防御性工作,而提供商层面的拒绝机制阻碍了这一过程。
In agentic coding, Mistral reports 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4.0, for a combined Artificial Analysis Coding Agent Index of 49.8%. Mistral notes these were evaluated privately ahead of the harness going public, so they are not yet independently reproducible.
在智能体编程方面,Mistral 报告其在 DeepSWE v1.1 上得分为 61.7%,在 SWE-Atlas-QnA 上得分为 59.4%,在 Terminal-Bench 4.0 上得分为 28.3%,综合 Artificial Analysis 编码智能体指数为 49.8%。Mistral 指出,这些结果是在基准测试工具公开之前私下评估的,因此目前尚无法独立复现。
A blind human evaluation run with Surge AI is the more honest signal. Professional annotators rated ML4 Preview 3.74 out of 5, second of 5 models, ahead of GLM-5.3 (3.60) and Kimi K3 (3.59), but behind Claude Opus 5 at 4.22.
由 Surge AI 执行的盲测人工评估是更可靠的信号。专业标注员对 ML4 Preview 的评分为 3.74/5,在 5 款模型中排名第二,优于 GLM-5.3(3.60)和 Kimi K3(3.59),但低于 Claude Opus 5 的 4.22。
On safety, ML4 resists 93.3% of attacks on Lakera’s B3 benchmark and scores 1.691 of a maximum 2.0 on KORABench.
在安全性方面,ML4 在 Lakera 的 B3 基准测试中抵御了 93.3% 的攻击,并在 KORABench 上获得 1.691 分(满分 2.0)。
How does it compare to its closest open-weight rivals?
它与最接近的开源权重竞争对手相比如何?
| Feature | Mistral Large 4 | DeepSeek V4 Pro | Kimi K3 | GLM-5.3 |
|---|---|---|---|---|
| Total parameters | 1.05T | 1.6T | 2.8T | Not officially published |
| Active per token | 49B | 49B | ~104B | Not officially published |
| Context window | 1M | 1M | 1M | 1M |
| Native image input | Yes | No | Yes | No |
| Weights available | Not yet, due end Oct 2026 | Yes, on Hugging Face | Yes, since 27 Jul 2026 | Yes, per Artificial Analysis |
| License | Not yet announced | MIT | Modified MIT | GLM-5.3 License |
| API price per 1M in/out | $1.36 / $4.18 | Varies by provider | $3.00 / $15.00 | $1.40 / $4.40 |
| Released | 6 Oct 2026 | Aug 2026 (0813 build) | 16 Jul 2026 | 14 Aug 2026 |
| 特性 | Mistral Large 4 | DeepSeek V4 Pro | Kimi K3 | GLM-5.3 |
|---|---|---|---|---|
| 总参数量 | 1.05T | 1.6T | 2.8T | 未官方公布 |
| 每 token 活跃参数 | 49B | 49B | ~104B | 未官方公布 |
| 上下文窗口 | 1M | 1M | 1M | 1M |
| 原生图像输入 | 是 | 否 | 是 | 否 |
| 权重可用性 | 尚未,预计 2026 年 10 月底结束 | 是,在 Hugging Face 上 | 是,自 2026 年 7 月 27 日起 | 是,根据 Artificial Analysis |
| 许可证 | 尚未宣布 | MIT | 修改版 MIT | GLM-5.3 许可证 |
| API 价格(每百万 in/out) | $1.36 / $4.18 | 因提供商而异 | $3.00 / $15.00 | $1.40 / $4.40 |
| 发布日期 | 2026 年 10 月 6 日 | 2026 年 8 月 (0813 构建) | 2026 年 7 月 16 日 | 2026 年 8 月 14 日 |
What can you build with it today?
今天你能用它构建什么?
The preview API supports function calling, structured outputs, document QnA, batching, and the Agents and Conversations endpoints. Cached input is priced at $0.14 per 1M tokens, which materially changes the economics of long-context agent loops at a 1M window.
预览版 API 支持函数调用、结构化输出、文档问答、批处理以及 Agents 和 Conversations 端点。缓存输入的定价为每百万 token 0.14 美元,这在 1M 上下文窗口下显著改变了长上下文智能体循环的经济性。
Key Takeaways
关键要点
- 1.05T total parameters, 49B active per token, 1M context, 1.6B vision encoder.
- Trained on 3,800 Grace Blackwell GPUs in Mistral’s own EU datacenters.
- API preview live now at $1.36 per 1M input and $4.18 per 1M output tokens.
- Weights promised by end of October 2026, so self hosting is not yet possible.
- Strongest results are in cybersecurity, where closed models often refuse the task.
- 1.05T 总参数量,每 token 49B 活跃参数,1M 上下文,1.6B 视觉编码器。
- 在 Mistral 自有的欧盟数据中心使用 3,800 块 Grace Blackwell GPU 进行训练。
- API 预览现已上线,输入价格为每百万 token 1.36 美元,输出价格为每百万 token 4.18 美元。
- 权重承诺于 2026 年 10 月底前提供,因此目前尚无法自行托管。
- 最强的结果出现在网络安全领域,而闭源模型通常在此类任务中会拒绝执行。
Check out the Mistral Large 4 announcement, Mistral Large 4 model docs and Artificial Analysis model comparisons. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
查看 Mistral Large 4 的公告、Mistral Large 4 模型文档以及 Artificial Analysis 的模型对比。所有功劳归于本项目的研究者。此外,欢迎在 Twitter 上关注我们,别忘了加入我们有 15 万+成员的 ML SubReddit 并订阅我们的新闻通讯。等等!你在 Telegram 上吗?现在你也可以加入我们的 Telegram 群组。
[Sponsored] The web is the one API most agents are missing. Databases, calendars and repos have APIs. The open web mostly doesn’t. The TinyFish MCP server gives any MCP client four tools: TinySearch, TinyFetch (full pages as markdown, JavaScript included), TinyBrowser for logins and forms, and TinyAgent for multi-step jobs. Search and Fetch are free.
[赞助] Web 是最缺少 API 的代理平台。数据库、日历和代码仓库都有 API。开放网络大多没有。TinyFish MCP 服务器为任何 MCP 客户端提供四种工具:TinySearch、TinyFetch(将完整页面转换为包含 JavaScript 的 Markdown)、用于登录和表单的 TinyBrowser,以及用于多步任务的 TinyAgent。Search 和 Fetch 功能免费。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力