跳到主内容
@wquguru
精选90Mistral AI News(RSS)模型发布/更新多源精选 ×9

Mistral发布开源多模态大模型Mistral Large 4

Introducing Mistral Large 4

原文
发到 X
推荐理由

全新旗舰级开源多模态模型发布,性能对标全球最强竞品,值得开发者关注。

Le chonk

胖乎乎

Introducing

隆重推出

Mistral Large 4

Back to Blog

返回博客

11 min read

11 分钟阅读

October 6, 2026

2026年10月6日

By Mistral

作者:Mistral

Share this post

分享此文章

Copy url to clipboardCopied

复制链接到剪贴板已复制

Le Chonk

胖乎乎

Today, we’re launching a public preview of Mistral Large 4. Unofficially ML4, very officially: le Chonk. ML4 pushes the frontier of open-weight performance. You can try the preview API today on Mistral Studio. Weights drop end of this month.

今天,我们推出了 Mistral Large 4 的公开预览。非官方简称 ML4,官方昵称:le Chonk(胖乎乎)。ML4 将开源权重模型的性能推向前沿。您今天即可在 Mistral Studio 上试用预览版 API。模型权重将于本月底发布。

Frontier performance

前沿性能

ML4 is a 1 trillion-parameter natively multimodal model with 49 billion active parameters. It is our largest and most capable model to date, and it continues to improve rapidly as we refine it.

ML4 是一个拥有 1 万亿参数、原生多模态的模型,其中活跃参数为 490 亿。这是我们迄今为止最大、能力最强的模型,随着我们不断优化,其性能仍在快速提升。

The model demonstrates exceptional performance across coding, agentic workflows, and multimodal understanding. It already achieves performance competitive with the strongest open-source models globally, while significantly outperforming any open-weight model developed in the US or Europe. On critical enterprise workloads, including cybersecurity, finance and law, we find it to be state-of-the-art among open models. In some domains such as visual grounding, it goes further still, surpassing even frontier closed models.

该模型在编码、智能体工作流和多模态理解方面展现出卓越的性能。它已达到与全球最强开源模型相媲美的水平,同时显著优于任何在美国或欧洲开发的开源权重模型。在网络安全、金融和法律等关键企业负载场景中,我们发现其处于开源模型的最先进水平。在某些领域,如视觉定位(visual grounding),它更进一步,甚至超越了领先的黑盒闭源模型。

We will release the weights by the end of the month. Until then, we are red-teaming the model in real-world settings with cybersecurity leaders, vetted partners, and state authorities, who will access the same model with reduced moderation and expanded cyber capabilities.

我们将在本月底前发布模型权重。在此之前,我们将与网络安全领域的领导者、经过审核的合作伙伴以及国家机构一起,在真实环境中对模型进行红队测试(red-teaming)。这些参与者将访问同一模型,但减少内容审核限制并扩展网络安全能力。

  • Coding - DeepSWE
  • Coding - Terminal Bench 4.0
  • Cyber
  • Agentic behaviour
  • Finance Agent
  • Harvey's Legal Agent
  • Grounding
  • 编码 - DeepSWE
  • 编码 - Terminal Bench 4.0
  • 网络安全
  • 智能体行为
  • 金融智能体
  • Harvey 法律智能体
  • 视觉定位

Forged in Europe. Built for AI sovereignty.

源自欧洲,为 AI 主权而建。

ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe. The public preview is served on that same infrastructure. It is a significant milestone in our long-term investment across infrastructure, research, and product development: state-of-the-art performance in critical verticals, delivered through open weights, designed to give customers control over their AI.

ML4 是在 Mistral 位于欧洲的自有数据中心中,使用 3,800 块 NVIDIA Grace Blackwell GPU 从头训练而成的。公开预览版也部署在同一基础设施上。这是我们在基础设施、研究和产品开发方面长期投资的重要里程碑:通过开源权重,在关键垂直领域提供最先进的性能,旨在让客户掌控自己的 AI。

This is particularly important in cybersecurity, where provider-level refusals can block legitimate vulnerability research and incident response, and where losing access to a capability mid-incident can itself become a critical security risk. ML4 pairs top-tier cyber performance with open weights and self-deployment, giving organizations both the capability and the autonomy to run advanced security work under their own policies.

这在网络安全领域尤为重要,因为服务提供商层面的拒绝可能会阻碍合法的漏洞研究和事件响应,而在事件处理过程中失去对某项能力的访问权限本身就会成为关键的安全风险。ML4 将顶级的网络性能与开放权重及自部署相结合,使组织既能获得能力,又能在自身策略下运行高级安全工作。

The model will be available across multiple regions worldwide, including a European deployment that Mistral operates end-to-end, independently of other digital service providers and under European law. Fun fact: a significant share of ML4’s training data was multilingual, spanning more than 160 languages, including every official language of the European Union.

该模型将在全球多个地区提供,包括由 Mistral 在欧洲全权运营、独立于其他数字服务提供商并遵循欧洲法律的欧洲部署。趣闻:ML4 的训练数据中有相当一部分是多语言的,涵盖超过 160 种语言,包括欧盟的所有官方语言。

We’ve been working closely with leading enterprises across the world in finance, engineering, manufacturing, logistics, pharmaceuticals, science, shipping, public sector, and other mission-critical industries to train ML4. In fact, the model uses the same training, customization, and RL environment we offer our customers through Mistral Forge.

我们一直与全球金融、工程、制造、物流、制药、科学、航运、公共部门及其他任务关键型行业的领先企业紧密合作,以训练 ML4。事实上,该模型使用了我们透过 Mistral Forge 向客户提供相同的训练、定制和强化学习环境。

Try it today

立即试用

There is still more to come. As we work toward releasing the weights, we will share further details on the model architecture, additional benchmarks, and our post-training methodology.

未来还有更多内容。在致力于发布权重的过程中,我们将分享有关模型架构、额外基准测试以及后训练方法的更多细节。

This model will also serve as the foundation for a new generation of specialized and optimized Mistral models. In the meantime, we invite you to try the preview API and share your feedback with us on social media.

该模型还将作为新一代专用且优化的 Mistral 模型的基础。与此同时,我们邀请您试用预览版 API,并通过社交媒体与我们分享反馈。

Capabilities deep-dive

能力深入解析

Cybersecurity

网络安全

ML4 is one of the world's strongest AI models for cybersecurity. On the Artificial Analysis Cyber Index, an independent evaluation of how well AI models find and fix security flaws in real software, it ranks among the top five models globally and leads open-weight models developed outside China by a wide margin. On one of the index's tests, which asks a model to reproduce a real vulnerability in open-source software and then patch it, ML4 scores 82%, the highest of any model. It also solves 93% of the challenges in Cybench, a set of 40 exercises drawn from security competitions, one of the highest scores reported for an open-weight model.

ML4 是全球最强的网络安全 AI 模型之一。在 Artificial Analysis Cyber Index(一项独立评估 AI 模型在真实软件中发现和修复安全缺陷能力的指标)中,它位列全球前五名,并以巨大优势领先于在中国境外开发的开放权重模型。在该指标的一项测试中,要求模型复现开源软件中的真实漏洞并进行修补,ML4 得分 82%,为所有模型中最高。此外,它在 Cybench(一组源自安全竞赛的 40 道练习题)中解决了 93% 的挑战,这是报告给开放权重模型的最高分数之一。

That top score reflects a practical advantage. Several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test because they refuse to perform the task. Yet defending software often starts with proving that a flaw is real, exactly the kind of work safety filters in closed models can block. This matters even more as threat actors increasingly jailbreak those same models to support offensive cyber activity: defenders need systems that can match those capabilities without being constrained by the same refusals. ML4 can do that work, and its capabilities extend beyond what it was explicitly trained for: in internal testing, it proved useful for analysing malware, prioritising vulnerabilities, and writing detection rules. For organisations that need sovereign, auditable AI for security operations, it will be able to run on private cloud or on-premise.

这一最高分反映了一个实际优势。包括 Claude Opus 5.5 和 GPT-6 Astra 在内的几款领先的闭源模型在同一测试中得分接近于零,因为它们拒绝执行该任务。然而,防御性软件通常始于证明漏洞确实存在,而这正是闭源模型中的安全过滤器可能阻止的工作类型。随着威胁行为者越来越多地越狱这些相同的模型以支持进攻性网络活动,这一点显得尤为重要:防御者需要能够匹配这些能力而不受相同拒绝限制的系统。ML4 可以完成这项工作,其能力还超出了它被明确训练的范围:在内部测试中,它被证明可用于分析恶意软件、优先处理漏洞以及编写检测规则。对于需要在安全运营中使用主权且可审计的 AI 的组织来说,它可以在私有云或本地部署上运行。

ML4 against the field : efficiently reasoning over diverse complex challenges

ML4 与业界对比:高效推理多样化的复杂挑战

Malware reverse-engineering: solving an out-of-distribution investigation task

恶意软件逆向工程:解决分布外调查任务

  • AA Cyber Index
  • CyberGym-E2E
  • Cybench

Agentic coding

智能体编码

ML4 excels across software engineering, repository understanding, and complex terminal workflows, scoring 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4. Its combined Coding Agent Index score of 49.8% places it ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max.

ML4 在软件工程、代码库理解和复杂终端工作流方面表现出色,在 DeepSWE v1.1 上得分为 61.7%,在 SWE-Atlas-QnA 上为 59.4%,在 Terminal-Bench 4 上为 28.3%。其综合 Coding Agent Index 得分为 49.8%,领先于 DeepSeek V4 Pro 0813 和 Qwen3.8 Max。

  • DeepSWE 1.1
  • Terminal Bench 4.0
  • SWE Atlas QnA

We also ran a blind human evaluation with Surge AI on coding quality: professional annotators rated model outputs on a 1–5 scale, with model identities hidden. ML4 Preview ranked second of five models (3.74), ahead of Kimi K3 (3.59), GLM-5.3 (3.60) and GLM-5.2 (3.40), and behind only Claude Opus 5 (4.22).

我们还与 Surge AI 进行了盲测人工评估,以评估编码质量:专业标注员对模型输出进行 1–5 分的评分,并隐藏了模型身份。ML4 Preview 在五个模型中排名第二(3.74 分),领先于 Kimi K3(3.59 分)、GLM-5.3(3.60 分)和 GLM-5.2(3.40 分),仅落后于 Claude Opus 5(4.22 分)。

Agentic Workflows

智能体工作流

ML4 runs general-purpose agents that gather information, use tools, and produce finished deliverables across complex workflows. On AutomationBench — 657 business workflows across apps like Gmail, Google Sheets, Slack, and Salesforce — it scores 59.9%, ahead of Kimi K3, MiMo-V2.6-Pro, and DeepSeek V4 Pro.

ML4 运行通用智能体,这些信息收集、使用工具并在复杂工作流中生成最终交付物。在 AutomationBench 上——涵盖 Gmail、Google Sheets、Slack 和 Salesforce 等应用的 657 个业务工作流——其得分为 59.9%,领先于 Kimi K3、MiMo-V2.6-Pro 和 DeepSeek V4 Pro。

It's just as strong on the professional deliverables that knowledge work actually produces: spreadsheets, slides, and PDFs. On AA-Briefcase, which evaluates long-horizon knowledge work, it reaches 1,393 Elo, ahead of DeepSeek V4 Pro.

它在知识工作实际产生的专业交付物上也同样强大:电子表格、演示文稿和 PDF。在评估长周期知识工作的 AA-Briefcase 上,它达到了 1,393 Elo 分,领先于 DeepSeek V4 Pro。

Multimodal

多模态

ML4 is a step change in the ability of our models to understand images. It reasons powerfully across complex documents, charts, and natural images, and brings vision to the industries where perception is critical such as engineering, manufacturing, and earth observation.

ML4 在模型理解图像的能力上实现了跨越式提升。它能够在复杂的文档、图表和自然图像中进行强大的推理,并将视觉能力带入对感知至关重要的行业,如工程、制造和地球观测。

The model can further combine visual grounding with agentic capabilities: from inspecting gigapixel satellite imagery — helping disaster-response teams act when time counts — to analyzing engineering-drawings — zooming in, inspecting, and verifying until the answer is exact. In our demos above, ML4 grounds dense natural scenes, verifies mechanical parts in technical drawings, retrieves evidence from PDFs, and scans massive geospatial images for the hardest-to-find objects.

该模型还能进一步将视觉定位与智能体(agentic)能力相结合:从检查吉像素级卫星图像——帮助灾害响应团队在关键时刻采取行动——到分析工程图纸——放大、检查并验证,直到答案精确无误。在上述演示中,ML4 能够定位密集的自然场景,验证技术图纸中的机械部件,从 PDF 中提取证据,并扫描大规模地理空间图像以寻找最难发现的物体。

On visual grounding particularly, we find ML4 to be one of the most capable models we tested, for instance surpassing GPT-6-Astra on Dense 200 (42% vs 41%).

特别是在视觉定位方面,我们发现 ML4 是我们测试过的最具能力的模型之一,例如在 Dense 200 基准测试中超越了 GPT-6-Astra(42% vs 41%)。

  • Dense 200
  • ChartQA Pro
  • GDP.pdf

Science and Math

ML4 brings strong scientific capabilities, built by combining AI-driven methods with our researchers' expertise in mathematics, physics, and chemistry.

ML4 带来了强大的科学能力,这是通过将 AI 驱动的方法与我们研究人员在数学、物理和化学领域的专业知识相结合而构建的。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →