跳到主内容
@wquguru
精选92Hacker News Best(web_list)模型发布/更新多源精选 ×8

Google发布Gemini 4 Argon模型,定价2美元/百万输入token

Gemini 4 Argon

原文
发到 X
推荐理由

全新旗舰模型发布,明确定价与百万级上下文能力,直接决定后续Agent开发成本与上限,从业者必读。

Gemini 4 Argon: our next era of frontier intelligence

Gemini 4 Argon:我们迈向前沿智能的新纪元

Sep 30, 2026

2026年9月30日

  • x.com
  • Facebook
  • LinkedIn
  • Mail
  • Copy link
  • x.com
  • Facebook
  • LinkedIn
  • 邮件
  • 复制链接

Gemini 4 Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.

Gemini 4 Argon 在现实世界中的复杂工作流中展现出前沿性能,涵盖软件工程、企业知识工作(如法律和金融)以及网络安全防御。

Koray Kavukcuoglu

SVP, Google DeepMind and Chief AI Architect, Google

Google DeepMind 高级副总裁兼 Google 首席 AI 架构师

Share

分享

  • x.com
  • Facebook
  • LinkedIn
  • Mail
  • Copy link
  • x.com
  • Facebook
  • LinkedIn
  • 邮件
  • 复制链接

Your browser does not support the audio element.

您的浏览器不支持音频元素。

Listen to article

收听文章

[[duration]] minutes

[[duration]] 分钟

This content is generated by Google AI. Generative AI is experimental

此内容由 Google AI 生成。生成式 AI 处于实验阶段。

Voice Speed

语音速度

Voice

语音

Speed 0.75X 1X 1.5X 2X

速度 0.75X 1X 1.5X 2X

Read AI-generated summary

阅读 AI 生成的摘要

  • Google’s new Gemini 4 Argon model brings advanced reasoning to complex, long-horizon professional tasks.
  • The model features an industry-leading 1 million token limit for deep, multi-step problem solving.
  • It excels at coding, financial research, legal drafting, and autonomous cybersecurity vulnerability patching.
  • Argon is currently rolling out to trusted cyber defenders through the Fairwind Program.
  • Google is prioritizing safety and rigorous testing before a wider release to the public.
  • 谷歌全新的 Gemini 4 Argon 模型将高级推理能力引入复杂、长周期的专业任务。
  • 该模型拥有行业领先的 100 万 token 限制,支持深度、多步骤的问题解决。
  • 它在编码、金融研究、法律起草和自主网络安全漏洞修补方面表现出色。
  • Argon 目前正通过 Fairwind 计划向受信任的网络防御者推出。
  • 在面向公众更广泛发布之前,谷歌优先考虑安全性和严格的测试。

Summaries were generated by Google AI. Generative AI is experimental.

摘要由 Google AI 生成。生成式 AI 处于实验阶段。

In this article

在本文中

Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program. Built to sustain deep reasoning across complex, long-horizon workflows, Argon is fundamentally changing the way we work and build at Google. It delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.

今天,我们宣布推出全新前沿模型 Gemini 4 Argon,该模型正通过我们的 Fairwind 计划向一组受信任的网络防御者推出。Argon 旨在跨复杂、长周期工作流维持深度推理,从根本上改变了我们在谷歌的工作和构建方式。它在现实世界的软件工程、企业知识工作(如法律和金融)以及网络安全防御等复杂工作流中提供了前沿性能。

Safely releasing frontier capabilities at this level requires a phased approach. We are actively engaged in the U.S. government’s voluntary process for pre-release model access while we gradually expand access. We’ll continue to gather feedback from early testers as we iterate on guardrails before making Argon available to developers, enterprises, and consumers as soon as possible.

安全地释放这一级别的前沿能力需要采取分阶段的方法。在我们逐步扩大访问权限的同时,我们积极参与美国政府关于发布前模型访问的自愿流程。随着我们在制定护栏规则并进行迭代,我们将继续收集早期测试者的反馈,以便尽快让开发人员、企业和消费者能够使用 Argon。

Argon will launch at an introductory price 1 of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price.

Argon 将以入门价格 1 推出,即每百万输入 token 2 美元,每百万输出 token 10 美元,缓存输入 token 的价格为输入 token 价格的 95% 折扣。

Changing how we work and build at Google

改变我们在谷歌的工作和构建方式

Gemini 4 Argon is already powering our internal workflows, with thousands of Googlers highlighting the model’s strengths in specialized coding tasks, conducting deeper research, and writing quality. It’s helping teams build faster and push the boundaries of engineering productivity and accelerating breakthroughs:

Gemini 4 Argon 已经在为我们的内部工作流提供动力,数千名谷歌员工强调了该模型在专业编码任务、进行更深入研究以及写作质量方面的优势。它正在帮助团队更快地构建产品,突破工程生产力的界限,并加速突破性进展:

  • Quantum algorithmic optimization: Argon is helping our quantum computing researchers optimize the spacetime resources (qubits × gates) of subroutines that bottleneck important applications. In one example, it beat the published baseline by 40% in a matter of minutes.
  • Memory efficiency: A team of Argon agents analyzed fleet-wide profiling telemetry to autonomously identify and apply memory optimizations across Google’s data centers, freeing up over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings.
  • Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google — scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia Zircon kernel. Given the criticality of many of these systems, such large-scale rewrites are undergoing rigorous automated and manual auditing, emulation testing, and review before rolling out to production.
  • For libgav1, Google's open source software for decoding video, Argon agents took an existing Rust port and replaced 32K lines of SIMD code by running many rounds of profile-guided experiments, studying the compiler's output, producing safe Rust so the compiler would vectorize it automatically. The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++.
  • 量子算法优化:Argon 正在帮助我们的量子计算研究人员优化对重要应用构成瓶颈的子例程的时空资源(量子比特 × 门)。在一个例子中,它在几分钟内就比已发布的基线高出 40%。
  • 内存效率:一组 Argon 代理分析了全车队的分析遥测数据,以自主识别并在谷歌的数据中心应用内存优化措施,一旦部署,将释放出超过 300 TiB 的内存,预计总节省量为 500 TiB 到 1 PiB。
  • 大规模代码库迁移与优化:Argon 代理正在 Google 范围内将 C/C++ 代码库迁移至 Rust——规模从 re2、libgav1 等核心库中的数万行,扩展到 Fuchsia Zircon 内核的 80 万行以上。鉴于许多此类系统的关键性,在部署到生产环境之前,这些大规模重写需经过严格的自动化和人工审计、模拟测试以及审查。
  • 对于 libgav1(Google 用于视频解码的开源软件),Argon 代理基于现有的 Rust 移植版本,通过运行多轮基于性能分析的实验、研究编译器输出,并生成安全 Rust 代码以使编译器能够自动进行向量化处理,从而替换了 3.2 万行 SIMD 代码。最终结果是一个内存安全的视频解码器,其运行速度比 Rust 移植版本快 2.7 倍,且视频输出完全一致,使其更接近优化后的 C++ 版本。

Working harder on your most complex problems

在你最复杂的问题上更努力地工作

To support Gemini 4 Argon’s capabilities across longer, more complex use cases, we are significantly expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens. When the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one go.

为支持 Gemini 4 Argon 在更长、更复杂用例中的能力,我们将模型的输出令牌限制大幅扩展至行业领先的 100 万令牌,此前为 6.4 万令牌。当模型有足够的空间进行深入思考并在单次轨迹中生成数十万令牌时,它便能在推理层面增加新的深度,从而一次性解决难题。

Enabling coding and enterprise workflows across domains

赋能跨领域的编码与企业工作流

Gemini 4 Argon’s capabilities across coding, reasoning, and multimodality and its ability to sustain long, multi-step tasks enable it to excel across a range of enterprise workflows.

Gemini 4 Argon 在编码、推理和多模态方面的能力,以及其维持长周期、多步骤任务的能力,使其能够在一系列企业工作流中脱颖而出。

Google engineers have been using Argon for their daily tasks, from everyday debugging to large-scale codebase migrations and algorithm designs. It sets a new state of the art on DeepSWE v1.1 (77.9%), which measures a model’s performance in real-world long-horizon software engineering tasks.

Google 工程师已将 Argon 用于日常任务,从日常调试到大规模代码库迁移和算法设计。它在 DeepSWE v1.1 上设定了新的最先进水平(77.9%),该基准衡量模型在真实世界长周期软件工程任务中的表现。

Beyond coding, Argon is the leading model on the Vals Index, which measures economic impact across finance, coding, legal, and tax work, with every sector weighted by its contribution to U.S. GDP. We see similarly leading performance across other domain specific evaluations, like Vals Finance Agent v2 (multi-step financial research) and Harvey’s Legal Agent Benchmark (legal research and drafting). On AutomationBench, Zapier’s benchmark measuring end-to-end execution across core business functions, Argon ranks #1 with a score of 51.3%.

除编码外,Argon 在 Vals Index 上也是领先模型,该指数衡量金融、编码、法律和税务工作的经济影响,每个领域均按其对美国 GDP 的贡献进行加权。在其他特定领域的评估中,我们也看到了同样领先的性能,例如 Vals Finance Agent v2(多步金融研究)和 Harvey’s Legal Agent Benchmark(法律研究与起草)。在 AutomationBench(Zapier 衡量核心业务功能端到端执行的基准)上,Argon 以 51.3% 的得分排名第一。

Argon is also uniquely strong when knowledge work requires visual understanding. It’s able to drive professional chart analysis, identify details from long videos, and take action based on a series of documents. For example, on LVBench, which measures long video understanding, Argon is state of the art with a score of 91.7%.

Argon 在需要视觉理解的知识型工作中同样具有独特优势。它能够驱动专业的图表分析,从长视频中识别细节,并基于一系列文档采取行动。例如,在衡量长视频理解能力的 LVBench 上,Argon 以 91.7% 的得分处于行业领先水平。

Leading in defensive cybersecurity

在防御性网络安全领域保持领先

To better equip cyber defenders for the new era of cyberattacks, we trained Gemini 4 Argon to be highly capable at cybersecurity defense. Argon can autonomously find, validate, and patch critical software vulnerabilities. For trusted defenders and our own internal teams at Google, we’ll be releasing Argon without cyber guardrails so they can leverage its full frontier-level cybersecurity defense capabilities.

为了更好地使网络防御者适应新的网络攻击时代,我们训练 Gemini 4 Argon 具备强大的网络安全防御能力。Argon 能够自主发现、验证并修补关键软件漏洞。对于受信任的防御者以及 Google 内部团队,我们将发布不带网络安全护栏的 Argon 版本,以便他们充分利用其前沿级别的网络安全防御能力。

Wiz is already using Argon for cybersecurity defense through its Scan for Good initiative – a program dedicated to protecting critical public infrastructure for free by finding and remediating high-risk exposures. In an early demonstration of its impact, the model uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide, identifying a severe risk that previous frontier models had missed.

Wiz 已通过其“Scan for Good”计划使用 Argon 进行网络安全防御——该计划致力于通过发现和修复高风险暴露来免费保护关键公共基础设施。在一项早期影响演示中,该模型发现了一个关键漏洞,导致全球医院使用的医疗软件中的敏感个人信息面临泄露风险,而此前其他前沿模型未能识别出这一严重风险。

On CWE-bench v1, which evaluates the model’s ability to remediate security vulnerabilities, Argon ties for first place with a top score of 68%, building on 3.8 Flash Cyber’s frontier performance on CWE-bench v0.

在评估模型修复安全漏洞能力的 CWE-bench v1 上,Argon 以 68% 的最高得分与第一名并列,这建立在 3.8 Flash Cyber 在 CWE-bench v0 上取得的前沿性能基础之上。

Gemini 4 Argon demonstrates impressive leaps in vulnerability discovery over 3.8 Flash Cyber. For example:

Gemini 4 Argon 在漏洞发现方面相比 3.8 Flash Cyber 展现出令人印象深刻的飞跃。例如:

  • On Google’s internal comprehensive vulnerability benchmark, Argon uncovered a wide range of exposures across complex codebases spanning 20 programming languages.
  • On Wiz’s internal black-box penetration testing benchmark, which tests a model’s ability to analyze live web systems without source code, Argon outperforms 3.8 Flash Cyber in discovering the attack surface, identifying vulnerabilities, and producing proof-of-concept evidence to validate them.
  • 在 Google 的内部综合漏洞基准测试中,Argon 在涵盖 20 种编程语言的复杂代码库中发现了一系列广泛的暴露点。
  • 在 Wiz 的内部黑盒渗透测试基准测试中(该测试评估模型在无源代码情况下分析实时 Web 系统的能力),Argon 在发现攻击面、识别漏洞以及生成概念验证证据以确认漏洞方面均优于 3.8 Flash Cyber。

Strengthening frontier safeguards before broad availability

在广泛可用之前加强前沿安全措施

Before rolling out Gemini 4 Argon broadly, we’re continuing to strengthen critical frontier safeguards across four main areas:

在全面推出 Gemini 4 Argon 之前,我们继续加强四个主要领域的关键前沿安全措施:

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →