Google DeepMind发布Gemini 4
Google DeepMind Unveils Gemini 4 Argon with 1M Output Tokens for Coding, Knowledge Work and Cyber Defense
百万级输出能力直接重塑Agent工作流上限,配合半价定价策略,对现有API格局构成实质性冲击,值得开发者重点关注。
Google DeepMind has just announced Gemini 4 Argon, its new frontier model and the first model of the Gemini 4 generation. It targets long-horizon software engineering, enterprise knowledge work in legal and finance, and cybersecurity defense. The biggest technical change is output length. Argon can generate up to 1M tokens in a single response, up from 64K on earlier Gemini models.
Google DeepMind 刚刚发布了 Gemini 4 Argon,这是其最新的前沿模型,也是 Gemini 4 世代的首个模型。它面向长周期软件工程、法律与金融领域的企业知识工作,以及网络安全防御。最大的技术变化在于输出长度。Argon 可在单次响应中生成多达 100 万 token,而此前的 Gemini 模型上限为 6.4 万 token。
What Google Announced
Google 宣布了什么
Google DeepMind described Argon as built for complex workflows across coding, enterprise knowledge work and cybersecurity defense.
Google DeepMind 将 Argon 描述为专为编码、企业知识工作和网络安全防御等复杂工作流而构建的模型。
Google is taking a phased approach. It is participating in the U.S. government’s voluntary process for pre-release model access. It will gather feedback from early testers and iterate on guardrails before a wider release.
Google 采取分阶段推进的方式。它正参与美国政府自愿性的发布前模型访问流程,将在更广泛发布之前收集早期测试者的反馈并迭代安全护栏(guardrails)。
Pricing is already public. Argon launches at an introductory $2 per 1M input tokens and $10 per 1M output tokens. Cached input tokens get a 95% discount, which works out to $0.10 per 1M. After the introductory period, pricing moves to $4 input and $20 output. Logan Kilpatrick confirmed the introductory $2 in and $10 out pricing.
定价已公开。Argon 以入门价每 100 万输入 token 2 美元、每 100 万输出 token 10 美元上线。缓存输入 token 享受 95% 折扣,折合每 100 万仅 0.10 美元。入门期结束后,价格将调整为每 100 万输入 4 美元、每 100 万输出 20 美元。Logan Kilpatrick 确认了入门期每 100 万输入 2 美元、每 100 万输出 10 美元的定价。
Why the 1M Output Limit Matters
为何 100 万输出限制至关重要
Current frontier APIs cap a single response far lower. Claude Opus 5.5, Claude Fable 5.1 and GPT-6 Astra each allow 128K output tokens.
当前前沿 API 的单次响应上限要低得多。Claude Opus 5.5、Claude Fable 5.1 和 GPT-6 Astra 均允许输出 12.8 万 token。
Google team states that Argon can think deeply and generate hundreds of thousands of tokens in one trajectory. For developers, that means large refactors or long reports without splitting work across turns. The cost is real, though. A full 1M output tokens costs $10 at introductory pricing and $20 after.
Google 团队表示,Argon 能够深入思考并在一次推理轨迹中生成数十万 token。对开发者而言,这意味着无需跨多轮对话拆分任务即可完成大型重构或撰写长篇报告。不过成本是实实在在的:按入门价计算,完整输出 100 万 token 需花费 10 美元,入门期后则为 20 美元。
Google has not disclosed Argon’s input context window.
Google 尚未披露 Argon 的输入上下文窗口大小。
Benchmarks: Where Argon Leads and Where It Trails
基准测试:Argon 的优势与劣势
Google compared Argon against GPT-6 Astra, Claude Opus 5.5 and Claude Fable 5.1. Argon leads outright on 12 of 18 benchmarks and ties for first on 1.
Google 将 Argon 与 GPT-6 Astra、Claude Opus 5.5 和 Claude Fable 5.1 进行了对比。在 18 项基准测试中,Argon 在 12 项上全面领先,另有 1 项并列第一。
Where it leads:
领先领域:
- DeepSWE v1.1 (long-horizon software engineering): 77.9%, a new state of the art. Opus 5.5 scores 74.2% and GPT-6 Astra 74.1%.
- Vals Index (economic impact across finance, coding, legal and tax): 68.9%, ranked first.
- AutomationBench (Zapier, end-to-end business execution): 51.3%, ranked first. Opus 5.5 scores 42.5%.
- Harvey Legal Agent Benchmark: 19.6%, against 5.4% for GPT-6 Astra.
- LVBench (long video understanding): 91.7%, a new state of the art.
- DeepSWE v1.1(长周期软件工程):77.9%,创下新的行业最佳水平(state of the art)。Opus 5.5 得分为 74.2%,GPT-6 Astra 为 74.1%。
- Vals Index(涵盖金融、编码、法律和税务的经济影响评估):68.9%,排名第一。
- AutomationBench(Zapier 端到端业务执行):51.3%,排名第一。Opus 5.5 得分为 42.5%。
- Harvey Legal Agent Benchmark(哈维法律代理基准测试):19.6%,而 GPT-6 Astra 仅为 5.4%。
- LVBench(长视频理解):91.7%,创下新的行业最佳水平。
Where it trails:
落后领域:
- FrontierSWE v2: 55.0%, behind GPT-6 Astra at 65.5%.
- Terminal-Bench 4.0: 57.4%, behind Claude Opus 5.5 at 66.4%.
- OSWorld-2.0 (computer use): 69.2%, behind GPT-6 Astra at 72.6%.
- FrontierSWE v2:55.0%,落后于 GPT-6 Astra 的 65.5%。
- Terminal-Bench 4.0:57.4%,落后于 Claude Opus 5.5 的 66.4%。
- OSWorld-2.0(计算机使用):69.2%,落后于 GPT-6 Astra 的 72.6%。
Artificial Analysis reported that Argon equals GPT-6 Astra on its Intelligence Index at 60% of the cost per task, using discounted prices.
Artificial Analysis 报告称,Argon 在智能指数上与 GPT-6 Astra 持平,且每项任务的成本仅为后者的 60%(采用折扣价格计算)。
Cyber Defense: Find, Validate, Patch
网络防御:发现、验证、修补
Google trained Argon to autonomously find, validate and patch critical software vulnerabilities. Trusted defenders and internal Google teams receive it without cyber guardrails.
Google 训练了 Argon 以自主发现、验证并修补关键软件漏洞。受信任的防御人员和 Google 内部团队可在无网络护栏限制的情况下使用它。
On CWE-bench v1, which tests vulnerability remediation, Argon ties for first at 68%. The rival models on that leaderboard run inside their own agent harnesses.
在测试漏洞修复能力的 CWE-bench v1 上,Argon 以 68% 的成绩并列第一。该排行榜上的其他模型均在各自的代理框架内运行。
Wiz is already using Argon through its Scan for Good initiative. The model found a critical vulnerability in healthcare software used by hospitals worldwide. Google says previous frontier models had missed it.
Wiz 已通过其“Scan for Good”计划开始使用 Argon。该模型发现了全球医院使用的医疗软件中的一个关键漏洞。Google 表示,此前的前沿模型均未能发现该漏洞。
Before broad release, Google is strengthening safeguards in 4 areas:
在广泛发布之前,Google 正在加强以下四个领域的安全保障措施:
- Misuse defenses for cyber and CBRN risks, including activation monitoring, under its Frontier Safety Framework.
- Indirect prompt injection resistance, where Argon leads Gray Swan’s IPI benchmark.
- Misalignment monitoring of chain-of-thought and actions, with the ability to stop execution.
- Sealed, isolated sandboxes for high-risk training and evaluations.
- 针对网络和 CBRN(化学、生物、放射性和核)风险的使用防范,包括在其前沿安全框架下的激活监控。
- 间接提示注入抵抗能力,Argon 在 Gray Swan 的 IPI 基准测试中领先。
- 对思维链和动作的对齐监控,具备停止执行的能力。
- 用于高风险训练和评估的密封隔离沙箱。
Argon Inside Google
Argon 在 Google 内部的应用
Thousands of Googlers already use Argon. Google shared 4 internal results:
数千名 Google 员工已在日常工作中使用 Argon。Google 分享了四项内部成果:
- Argon agents applied memory optimizations across data centers, freeing over 300 TiB, with 500 TiB to 1 PiB projected.
- Agents replaced 32K lines of SIMD code in the libgav1 Rust port. The decoder runs 2.7x faster with identical output.
- Agents are migrating C/C++ codebases to Rust, up to 800K+ lines in the Fuchsia Zircon kernel.
- Argon beat a published quantum algorithm baseline by 40% in minutes.
- Argon 代理在数据中心实施了内存优化,释放了超过 300 TiB 的空间,预计未来可释放 500 TiB 至 1 PiB。
- 代理替换了 libgav1 Rust 移植版中的 32K 行 SIMD 代码。解码器的运行速度提升了 2.7 倍,且输出结果完全一致。
- 代理正在将 C/C++ 代码库迁移至 Rust,涉及 Fuchsia Zircon 内核中多达 800K+ 行的代码。
- Argon 在几分钟内就比已发布的量子算法基线高出 40%。
Comparison: Gemini 4 Argon vs Closest Competitors
对比:Gemini 4 Argon 与最接近的竞争对手
| Feature | Gemini 4 Argon | Claude Opus 5.5 | Claude Fable 5.1 | GPT-6 Astra |
|---|---|---|---|---|
| Developer | Google DeepMind | Anthropic | Anthropic | OpenAI |
| Availability | Fairwind Program only | Claude API and clouds | Claude API and clouds | OpenAI API |
| Max output per response | 1M tokens | 128K | 128K | 128K |
| Context window | Not disclosed | 1M | 1M | 1.05M |
| Input / output price (per 1M) | $2 / $10 intro, then $4 / $20 | $4 / $20 | $10 / $50 | $10 / $50 |
| Cached input (per 1M) | $0.10 (intro) | $0.20 | $0.25 | $1.00 |
| Open weights | No | No | No | No |
| DeepSWE v1.1 | 77.9% | 74.2% | 67.4% | 74.1% |
| Vals Index | 68.9% | 67.0% | 65.8% | 63.1% |
| FrontierSWE v2 | 55.0% | 62.3% | 56.3% | 65.5% |
| Terminal-Bench 4.0 | 57.4% | 66.4% | 57.9% | 58.2% |
| CWE-bench v1 | 68% (tie) | 67% | 58% | 68% (tie) |
| 功能 | Gemini 4 Argon | Claude Opus 5.5 | Claude Fable 5.1 | GPT-6 Astra |
|---|---|---|---|---|
| 开发者 | Google DeepMind | Anthropic | Anthropic | OpenAI |
| 可用性 | 仅限 Fairwind 计划 | Claude API 及云平台 | Claude API 及云平台 | OpenAI API |
| 单次响应最大输出 | 1M tokens | 128K | 128K | 128K |
| 上下文窗口 | 未披露 | 1M | 1M | 1.05M |
| 输入/输出价格(每 1M tokens) | $2 / $10 促销价,之后为 $4 / $20 | $4 / $20 | $10 / $50 | $10 / $50 |
| 缓存输入(每 1M tokens) | $0.10(促销价) | $0.20 | $0.25 | $1.00 |
| 开放权重 | 否 | 否 | 否 | 否 |
| DeepSWE v1.1 | 77.9% | 74.2% | 67.4% | 74.1% |
| Vals Index | 68.9% | 67.0% | 65.8% | 63.1% |
| FrontierSWE v2 | 55.0% | 62.3% | 56.3% | 65.5% |
| Terminal-Bench 4.0 | 57.4% | 66.4% | 57.9% | 58.2% |
| CWE-bench v1 | 68%(并列) | 67% | 58% | 68%(并列) |
Sources: Google, Anthropic Opus pricing, Anthropic Fable 5.1 docs, OpenAI GPT-6 Astra docs, OpenRouter. Benchmark scores are from Google’s published comparison. GPT-6 Astra prices are its short-context tier.
来源:Google、Anthropic Opus 定价、Anthropic Fable 5.1 文档、OpenAI GPT-6 Astra 文档、OpenRouter。基准测试分数来自 Google 发布的对比数据。GPT-6 Astra 的价格为其短上下文层级价格。
Key Takeaways
关键要点
- Gemini 4 Argon raises the output limit from 64K to 1M tokens.
- It leads DeepSWE v1.1 (77.9%) and the Vals Index (68.9%).
- It trails on FrontierSWE v2, Terminal-Bench 4.0 and OSWorld-2.0.
- Introductory pricing of $2 / $10 is half of Claude Opus 5.5.
- Access is limited to Fairwind cyber defenders; no public release date yet.
- Gemini 4 Argon 将输出限制从 64K 提升至 1M tokens。
- 它在 DeepSWE v1.1(77.9%)和 Vals Index(68.9%)上领先。
- 它在 FrontierSWE v2、Terminal-Bench 4.0 和 OSWorld-2.0 上落后。
- 入门定价为 $2 / $10,仅为 Claude Opus 5.5 的一半。
- 仅限 Fairwind 网络防御者访问;尚未公布公开发布日期。
Check out the technical details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
查看技术细节。所有功劳归于本项目的研究者。也欢迎在 Twitter 上关注我们,别忘了加入我们拥有 150k+ 成员的 ML SubReddit 并订阅我们的新闻通讯。等等!你在 Telegram 上吗?现在你也可以在 Telegram 上加入我们。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力