跳到主内容
@wquguru
精选88Latent Space(RSS)模型发布/更新多源精选 ×12

Google DeepMind 发布 Gemini 4

[AINews] Gemini 4 Argon: GDM’s answer to Astra/Fable, with 1M output

原文
发到 X
推荐理由

Gemini 4 Argon 是 Google 重返前沿的重大旗舰发布,百万级输出能力极具看点,建议关注其实际落地效果与定价策略。

GDM last shipped a larger-than-Flash model in February (3.1 Pro), and after successive incremental 3.x Flash versions and the big GDM management shakeup last month, the largest question for GDM was when they would catch up to peers who have in the meantime launched Fable and Astra class models.

GDM 在二月(3.1 Pro)交付了一个比 Flash 更大的模型,而在经历了连续的增量 3.x Flash 版本以及上个月 GDM 管理层的重大重组后,GDM 面临的最大问题是他们何时能追上在此期间已推出 Fable 和 Astra 级模型的同行。

Well, Argon’s here, with VERY respectable benchmarks (SOTA in 13 of 19 credible benchmarks)… but only accessible in limited cybersecurity preview, though access is promised “as soon as possible”:

好吧,Argon 来了,拥有非常令人尊敬的基准测试结果(在 19 项可信基准测试中 13 项处于 SOTA 水平)……但目前仅在有限的网络安全预览版中可用,不过承诺将“尽快”开放访问:

We like the experimental Long Decode Continuation, which increases output tokens up to 1M as an industry first.

我们喜欢实验性的 Long Decode Continuation(长解码延续)功能,它将输出 token 数量提升至 1M,这是行业首创。

AI News for 9/29/2026-9/30/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!

2026年9月29日至9月30日的 AI 新闻。我们检查了 12 个 subreddit、544 条推文,没有进一步的 Discord 讨论。AINews 的网站允许你搜索所有过往期号。提醒一下,AINews 现在是 Latent Space 的一个板块。你可以选择订阅或退订电子邮件频率!

AI Twitter Recap

AI Twitter 回顾

Gemini 4 Argon: Google Returns to the Frontier

Gemini 4 Argon:Google 重返前沿

  • Launch: Google DeepMind introduced Gemini 4 Argon for coding, enterprise knowledge work and cyber defense (@GoogleDeepMind, @sundarpichai).
  • Availability: Access starts with government users and trusted cyber defenders in the Fairwind Program. Google says it will refine guardrails before opening access to developers, enterprises and consumers (@Google, @demishassabis).
  • Output limit: Google cites an industry-leading 1M-token output limit, up from 64K (@GoogleAI, @TheRundownAI).
  • Measurement note: Vals lists 262K max output. Artificial Analysis reached 1M output tokens through Long Decode Continuation, a new API feature that pauses long responses and resumes them across calls (@ValsAI, @ArtificialAnlys).
  • Pricing: Standard pricing is $4/$20 per 1M input/output tokens. A 50% introductory discount brings it to $2/$10, with no end date announced. Cached input gets a 95% discount (@_philschmid, @ArtificialAnlys).
  • Google’s claimed results: Argon takes first place on 13 of 19 published benchmarks against GPT-6 Astra and Claude Opus 5.5. On DeepSWE it scores 77.9%, versus 74.2% for Opus 5.5 and 74.1% for Astra (@TheRundownAI).
  • Internal deployments: Google reports that Argon agents freed more than 300 TiB of data-center memory and are migrating more than 800K lines of C/C++ kernel code to Rust (@kimmonismus).
  • Video decoder: Agents replaced 32K lines of SIMD code with safe Rust, making the existing Rust port 2.7x faster with identical output.
  • Research use: The team says internal agent loops built on Argon helped complete the CK conjecture (@mirrokni).
  • Artificial Analysis evaluation: Argon scores 53 on the Intelligence Index, matching GPT-6 Astra (53) and edging GPT-6.1 Sol (52) (@ArtificialAnlys).
  • Cost per task: At discounted pricing it costs $1.99 per task versus $3.26 for Astra; standard pricing would raise this to $3.98.
  • Token use: The savings come from price, not efficiency. Argon averages 62K output tokens per task against Astra’s 27K.
  • Agentic work: It ranks #1 on AutomationBench-AA at 77.5% and scores 57% on Terminal Bench 4, behind Sonnet 5.5, Opus 5.5 and Astra.
  • Hallucination: Its 15% rate on AA-Omniscience compares with 51% for Astra. The tradeoff is lower accuracy: 50% versus Astra’s 63% (@aipulseda1ly).
  • Vals evaluation: Argon is #1 on the Vals Index at 68.9%, at an average $15.68 per task (@ValsAI, @ValsAI).
  • Coding: It built 30 Vibe Code Bench apps perfectly, against 25 for Opus 5 and 24 for Astra (@ValsAI).
  • Terminal and security: Terminal-Bench 4.0 rose from 19.0% to 57.6%. It scores 70% on CyberBench proof-of-concept tasks and 100% on IOI 2024–2026 (@ValsAI).
  • Efficiency: It uses about a quarter of Sonnet 5.5’s output tokens on Vals Index tasks (@ValsAI).
  • Arena and other evals: Argon is #1 in Text Arena at 1525 and #8 in Code Arena WebDev at 1679 (@arena).
  • Agent Arena: It ranks #8 overall and #1 for steerability on a preliminary 3K sessions (@arena).
  • PostTrainBench: It scores 45.3%, up from 21.99% for Gemini 3.1 Pro (@karinanguyen).
  • Skepticism: Some observers questioned the published numbers.
  • Legal benchmark: Argon’s reported 19.6% on Harvey’s legal benchmark trails Muse Spark 1.2’s listed 25.42% (@BlackHC).
  • Other critiques: Commentators raised possible preference-data benchmaxxing and objected to some figures, including DeepSWE (@teortaxesTex, @teortaxesTex).
  • 发布:Google DeepMind 推出了 Gemini 4 Argon,用于编码、企业知识工作和网络防御 (@GoogleDeepMind, @sundarpichai)。
  • 可用性:访问权限首先向政府用户和 Fairwind 计划中的受信任网络防御者开放。Google 表示将在向开发者、企业和消费者开放访问之前完善护栏机制 (@Google, @demishassabis)。
  • 输出限制:Google 引用了行业领先的 1M token 输出限制,此前为 64K (@GoogleAI, @TheRundownAI)。
  • 测量说明:Vals 列出最大输出为 262K。Artificial Analysis 通过 Long Decode Continuation(长解码延续)达到了 1M 输出 token,这是一个新的 API 功能,可暂停长响应并在多次调用中恢复 (@ValsAI, @ArtificialAnlys)。
  • 定价:标准价格为每 1M 输入/输出 token 4美元/20美元。50% 的入门折扣将其降至 2美元/10美元,未公布结束日期。缓存输入可获得 95% 的折扣 (@_philschmid, @ArtificialAnlys)。
  • Google 声称的结果:在与 GPT-6 Astra 和 Claude Opus 5.5 对比的 19 项已发布基准测试中,Argon 在 13 项中排名第一。在 DeepSWE 上,其得分为 77.9%,而 Opus 5.5 为 74.2%,Astra 为 74.1% (@TheRundownAI)。
  • 内部部署:Google 报告称,Argon 代理释放了超过 300 TiB 的数据中心内存,并正在将超过 80万行 C/C++ 内核代码迁移到 Rust (@kimmonismus)。
  • 视频解码器:代理用安全的 Rust 代码替换了 32K 行的 SIMD 代码,使现有的 Rust 端口速度提高了 2.7 倍,且输出结果相同。
  • 研究用途:团队表示,基于 Argon 构建的内部代理循环帮助完成了 CK 猜想 (@mirrokni)。
  • Artificial Analysis 评估:Argon 在 Intelligence Index 上得分为 53,与 GPT-6 Astra(53)持平,并略高于 GPT-6.1 Sol(52)(@ArtificialAnlys)。
  • 每任务成本:折扣定价下,每任务成本为 1.99 美元,而 Astra 为 3.26 美元;标准定价会将此成本提高至 3.98 美元。
  • Token 使用量:节省来自价格而非效率。Argon 平均每任务输出 62K tokens,而 Astra 为 27K。
  • 智能体工作:它在 AutomationBench-AA 上以 77.5% 的得分排名第一,在 Terminal Bench 4 上得分为 57%,落后于 Sonnet 5.5、Opus 5.5 和 Astra。
  • 幻觉:其在 AA-Omniscience 上的幻觉率为 15%,而 Astra 为 51%。代价是准确率较低:50% 对比 Astra 的 63% (@aipulseda1ly)。
  • Vals 评估:Argon 在 Vals Index 上以 68.9% 的得分排名第一,平均每任务成本为 15.68 美元 (@ValsAI, @ValsAI)。
  • 编程:它在 Vibe Code Bench 中完美构建了 30 个应用,而 Opus 5 为 25 个,Astra 为 24 个 (@ValsAI)。
  • 终端与安全:Terminal-Bench 4.0 的得分从 19.0% 提升至 57.6%。它在 CyberBench 概念验证任务上得分为 70%,在 IOI 2024–2026 上得分为 100% (@ValsAI)。
  • 效率:在 Vals Index 任务中,它使用的输出 Token 数量约为 Sonnet 5.5 的四分之一 (@ValsAI)。
  • Arena 及其他评估:Argon 在 Text Arena 中以 1525 分排名第一,在 Code Arena WebDev 中以 1679 分排名第八 (@arena)。
  • Agent Arena:在初步的 3K 会话中,它总体排名第八,在可控性方面排名第一 (@arena)。
  • PostTrainBench:其得分为 45.3%,高于 Gemini 3.1 Pro 的 21.99% (@karinanguyen)。
  • 质疑:一些观察家对公布的数字表示怀疑。
  • 法律基准测试:Argon 在 Harvey 法律基准测试中报告的 19.6% 低于 Muse Spark 1.2 列出的 25.42% (@BlackHC)。
  • 其他批评:评论者提出了可能存在偏好数据刷榜的问题,并对部分数据提出异议,包括 DeepSWE (@teortaxesTex, @teortaxesTex)。

GPT-6.1 Sol and OpenAI’s DevDay Agent Stack

GPT-6.1 Sol 及 OpenAI 的 DevDay Agent Stack

  • Independent evals: GPT-6.1 Sol is the new #1 on MathArena (@j_dekoninck).
  • Code Arena: It ranks #3 on WebDev at 1759, 70 points above GPT-6 Sol for the same $2/$10 pricing (@arena).
  • Cost per task: Artificial Analysis measures $0.72 per task at max effort, versus $3.26 for Astra and $1.04 for GPT-6 Sol (@ArtificialAnlys).
  • Source of savings: Sol uses fewer turns and has a lower cache-read price (@ArtificialAnlys).
  • Luna bug fix: OpenAI fixed an image-encoding bug, adding 1 Intelligence Index point to GPT-6 Luna.
  • Ultrafast inference: OpenAI quotes up to 300 tok/s. SemiAnalysis reports it runs on NVIDIA GPUs at low batch sizes, not on Cerebras (@kimmonismus).
  • Hands-on report: Generation is about 8x faster, but end-to-end agent tasks speed up only 2–4x because tool latency dominates (@sayashk).
  • Computer use: Gains are largest here, since UI actions respond in milliseconds.
  • Cost: The tester exhausted a weekly limit in about 2 hours.
  • Product layer: DevDay introduced dots (persistent agents with their own cloud computers), a Decisions API and computer use (@latentspacepod).
  • Sites: ChatGPT Sites can now host MCP servers and turn them into installable plugins (@mxstbr).
  • Usage limits: Users report one-off credits worth about $2,500. Others complain that usage limits were cut (@kimmonismus, @kimmonismus).
  • 独立评估:GPT-6.1 Sol 成为 MathArena 的新冠军 (@j_dekoninck)。
  • Code Arena:它在 WebDev 上以 1759 分排名第三,比相同定价(2/10 美元)下的 GPT-6 Sol 高出 70 分 (@arena)。
  • 每任务成本:Artificial Analysis 测得最大努力下的每任务成本为 0.72 美元,而 Astra 为 3.26 美元,GPT-6 Sol 为 1.04 美元 (@ArtificialAnlys)。
  • 节省来源:Sol 使用的轮次更少,且缓存读取价格更低 (@ArtificialAnlys)。
  • Luna 漏洞修复:OpenAI 修复了一个图像编码漏洞,使 GPT-6 Luna 的 Intelligence Index 得分增加 1 点。
  • 超快推理:OpenAI 称其速度高达每秒 300 个 token。SemiAnalysis 报道称它在 NVIDIA GPU 上以低批次大小运行,而非在 Cerebras 上运行 (@kimmonismus)。
  • 实测报告:Generation 速度提升约 8 倍,但端到端智能体任务仅加速 2–4 倍,因为工具延迟占主导 (@sayashk)。
  • 计算机使用:此处增益最大,因为 UI 操作响应在毫秒级。
  • 成本:测试人员在约 2 小时内耗尽了每周额度限制。
  • 产品层:DevDay 推出了 dots(拥有独立云电脑的持久化智能体)、Decisions API 和计算机使用功能 (@latentspacepod)。
  • 站点:ChatGPT Sites 现在可以托管 MCP 服务器并将其转化为可安装的插件 (@mxstbr)。
  • 使用限制:用户报告称获得了一次性赠款,价值约 2,500 美元。其他人抱怨使用限额被削减 (@kimmonismus, @kimmonismus)。

Other Releases: Embeddings, Image/Video and Open Models

其他发布:Embeddings、图像/视频和开放模型

  • Perplexity contextual embeddings: pplx-embed-v2-context-9b-preview is open on Hugging Face (@perplexity_ai).
  • Method: The model encodes the whole document once and pools chunk vectors afterward. Training distills relevance from a context-compression model instead of using single gold-chunk labels (@denisyarats).
  • Results: It sets a new state of the art on ConTEB. On turbopuffer’s private context-bench it beats voyage-context-4 by 14.4 points in answer recall@10, using 1 KB int8 vectors against 8 KB (@turbopuffer).
  • Cohere Embed 5: The family has Pro and Fast variants in a shared embedding space, so you can index with one and retrieve with the other (@cohere).
  • Fast tier: Cohere says it beats other fast-tier models by at least 6 points at a third less cost than Pro. Evaluation uses its new RCP-nDCG@10 metric (@cohere).
  • Ideogram 4.5: The editing model targets artifact-free multi-turn edits, with open weights promised (@ideogram_ai).
  • Edit fidelity: Over ten consecutive edits, 94–99% of untouched content stays identical (@fal).
  • Ranking: It is #18 in Image Edit Arena at 1351 (@arena).
  • Video benchmark: Artificial Analysis launched AA-Video-T2V v2.0, judged at 1080p with more than 68K human votes (@ArtificialAnlys).
  • Leaders: Wan 3.0 is #1 at $12/min. Seedance 2.5 is #2 at $34.12/min, and MiniMax H3 is statistically tied at $4.80/min.
  • Utopai X: This post-train of MiniMax H3 debuts at #2 (@ArtificialAnlys).
  • Open and small models:
  • Ling-3.1-flash: A 500B model reported close to GPT-5.6 Sol and Opus 5 (@kimmonismus). It ranks #2 among open-weight models in Mobile App Arena (@DesignArena).
  • Praxis-1: Runway released an open-weight world-action model and says robotics policy performance scales predictably with third-person video (@agermanidis).
  • Solar Mini 4: Upstage reports 35B total / 3B active parameters. It scores 24 on the Intelligence Index at $0.10/$0.40 (@ArtificialAnlys).
  • Caching penalty: It still costs about 5x Luna per task, because only 48% of its repeated context hits cache versus 99% for Luna (@ArtificialAnlys).
  • Perplexity 上下文嵌入:pplx-embed-v2-context-9b-preview 已在 Hugging Face 开源 (@perplexity_ai)。
  • 方法:该模型对整篇文档进行一次编码,随后对块向量进行池化。训练过程从上下文压缩模型中蒸馏相关性,而非使用单个黄金块标签 (@denisyarats)。
  • 结果:它在 ConTEB 上创下新的最先进水平。在 turbopuffer 的私有 context-bench 上,它使用 1 KB int8 向量(对比 8 KB),在 answer recall@10 指标上比 voyage-context-4 高出 14.4 分 (@turbopuffer)。
  • Cohere Embed 5:该系列包含 Pro 和 Fast 变体,共享同一嵌入空间,因此你可以用一种变体建立索引,用另一种进行检索 (@cohere)。
  • 快速层级:Cohere 表示其快速层级模型性能比其他同类模型高出至少 6 分,且成本仅为 Pro 版本的三分之一。评估使用了其新的 RCP-nDCG@10 指标 (@cohere)。
  • Ideogram 4.5:该编辑模型旨在实现无伪影的多轮编辑,并承诺开源权重 (@ideogram_ai)。
  • 编辑保真度:在连续十次编辑中,94–99% 的未修改内容保持完全一致 (@fal)。
  • 排名:它在 Image Edit Arena 中排名第 18 位,得分为 1351 (@arena)。
  • 视频基准测试:Artificial Analysis 推出了 AA-Video-T2V v2.0,以 1080p 分辨率评判,并获得超过 68K 张人工投票 (@ArtificialAnlys)。
  • 领先者:Wan 3.0 以每分钟 12 美元的价格排名第一。Seedance 2.5 以每分钟 34.12 美元排名第二,MiniMax H3 以每分钟 4.80 美元的价格在统计上与第二名持平。
  • Utopai X:这是 MiniMax H3 的后训练版本,首发即排名第二 (@ArtificialAnlys)。
  • 开放和小模型:
  • Ling-3.1-flash:一款 500B 参数模型,据报道性能接近 GPT-5.6 Sol 和 Opus 5 (@kimmonismus)。它在 Mobile App Arena 的开放权重模型中排名第二 (@DesignArena)。
  • Praxis-1:Runway 发布了一款开放权重的世界动作模型,并表示机器人策略性能与第三人称视频呈可预测的缩放关系 (@agermanidis)。
  • Solar Mini 4:Upstage 报告总参数量为 35B,活跃参数量为 3B。其 Intelligence Index 得分为 24,价格为 $0.10/$0.40(@ArtificialAnlys)。
  • 缓存惩罚:由于仅 48% 的重复上下文命中缓存,而 Luna 为 99%,因此每个任务的成本仍约为 Luna 的 5 倍(@ArtificialAnlys)。

Agent Research, Inference and Systems

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

关联信息,但可能不是同一事件