跳到主内容
精选80Latent Space(RSS)模型发布/更新多源精选 ×4

Grok 4.6发布,xAI称其为知识工作第二强模型

[AINews] SpaceXAI Grok 4.6 and Grok @Bot

原文
推荐理由

做Agent和知识工作场景的同学注意了,Grok 4.6在价格/性能上很有竞争力,且强调长时Agent能力,建议对比测试你的工作负载。

One of our top recurring themes of the year has been coding agents breaking containment into knowledge work, and it’s clear that the AI teammate/multiplayer/multiagent space is the next big AI battleground. With Claude Tag launching to mixed reviews and Block’s Buzz requiring a more technical user, the space was still open for a new category leader, which the now Cursor→SpaceX team has adroitly shipped to very positive reviews:

我们今年反复出现的首要主题之一是编码代理突破界限进入知识工作领域,很明显,AI队友/多人/多代理空间是下一个重大的AI战场。随着Claude Tag发布后评价褒贬不一,Block的Buzz需要更技术性的用户,这个领域仍然为新的类别领导者敞开着大门,而现在Cursor→SpaceX团队已经巧妙地推出了产品,并获得了非常积极的评价:

This is powered by their newest model, Grok 4.6, released today as arguably the second best knowledge work model in the world (as both competitor Cognition and Elon acknowledges)… though it is surely the top by efficiency:

这得益于他们最新的模型Grok 4.6,今天发布,可以说是世界上第二好的知识工作模型(正如竞争对手Cognition和Elon所承认的那样)……尽管在效率方面它无疑是最顶尖的:

Grok 4.6 is a confirmed 1.5T model that “builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work”. The only training disclosure can be reproduced in full (emphasis ours):

Grok 4.6被确认为1.5T模型,它“在Grok 4.5的基础上构建,特别关注长时间运行的代理和更雄心勃勃的交互式及视觉工作”。唯一的训练披露可以完整复现如下(强调为我们所加):

Grok 4.6 underwent a longer supplemental training run than Grok 4.5, with curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe. This produced a stronger foundation for the SFT and RL stages that followed.

Grok 4.6经历了比Grok 4.5更长的补充训练运行,使用了精选的模型生成数据,用于推理和高级技术概念,高质量工程数据,以及改进的优化器和训练配方。这为后续的SFT和RL阶段奠定了更坚实的基础。

We then used Grok 4.5 to regenerate the SFT trajectories across reasoning efforts, agent harnesses, and domains such as STEM, software engineering, and knowledge work, and filtered out problematic traces with model-based checks. The resulting SFT checkpoint shows strong performance and improved behavior.

然后我们使用Grok 4.5重新生成了跨推理努力、代理工具以及STEM、软件工程和知识工作等领域的SFT轨迹,并通过基于模型的检查过滤掉了有问题的轨迹。由此产生的SFT检查点表现出强大的性能和改进的行为。

Grok 4.6 is trained on a wide range of agentic RL tasks, including knowledge work, general coding, and domain-specific environments for kernel optimization, web development, computer-aided design, and more.

Grok 4.6在广泛的代理RL任务上进行了训练,包括知识工作、通用编码以及针对内核优化、Web开发、计算机辅助设计等领域的特定环境。

It is currently unclear if the lack of sandbox escape incidents when it came to training Grok 4.6 is a testament to their infra engineers or an indictment of their researchers.

目前尚不清楚在训练Grok 4.6时缺乏沙箱逃逸事件是证明了他们的基础设施工程师的能力,还是对研究人员的一种控诉。

(that is a joke about current events, don’t get mad)

(这是对当前事件的一个玩笑,别生气)

AI News for 8/11/2026-8/12/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!

2026年8月11日至8月12日的AI新闻。我们检查了12个子版块、544条推特,没有进一步的Discord。AINews的网站允许你搜索所有过去的期刊。提醒一下,AINews现在是Latent Space的一个板块。你可以选择接收或不接收电子邮件频率!

AI Twitter Recap

AI推特回顾

Frontier Model Day: Grok 4.6, Qwen3.8-Max, DeepSeek V4 Pro, and Microsoft’s MAI-Thinking-1

前沿模型日:Grok 4.6、Qwen3.8-Max、DeepSeek V4 Pro和微软的MAI-Thinking-1

  • Grok 4.6 reaches the frontier on price/performance: xAI released Grok 4.6, described as a major step up from 4.5 at the same price. Independent evaluations from Artificial Analysis place it at 61 on the Intelligence Index, roughly in line with GPT-5.6 Sol Max, behind Claude Opus/Fable, with strong agentic results including 88.4% on Terminal-Bench v2.1, 1753 GDPval-AA v2 Elo, and competitive AA-Briefcase performance at far lower cost (AA-Briefcase note). Early arena data from Code Arena also slots it near GPT-5.6 Sol and Claude Fable on webdev tasks. Pricing is a central theme: AA highlights $2/$6 per 1M input/output tokens, materially below frontier peers, while practitioners immediately framed it as the new default for coding and bug-finding workloads (Pawel Huryn, Cognition availability in Devin). xAI says the gains came from a longer supplemental training run, regenerated SFT traces, and agentic RL over coding, web, CAD, and kernel optimization; they also report more self-testing behavior during long tasks (@kimmonismus summary). Elon also said Grok 4.7 is already in flight, with initial training complete and supplemental training on SpaceX internal data planned.
  • Qwen3.8-Max open weights are out: Alibaba’s Qwen3.8-Max dropped as an open-weight 2.4T total / 95B active MoE. Community notes emphasize its scale, day-0 serving, and long-context/agent orientation: Yuchen Jin called it one of the largest open-weight releases to date; vLLM shipped day-0 support plus vendor-specific 4-bit checkpoints for NVIDIA B300 and AMD MI355X; Together AI and Baseten also announced immediate support. One important caveat from users: the released open-weights variant appears to be text-only, with no vision input in the initial drop (skalskip92).
  • DeepSeek V4 Pro GA undercuts the market: DeepSeek’s V4 Pro GA rollout immediately drew attention less for “best benchmark in every column” than for economics. Multiple observers highlighted pricing around $0.435/M input and $0.87/M output (kimmonismus), with Cline calling it roughly 57× cheaper than Fable 5 while reporting meaningful gains over the preview, including a 15.8% Terminal Bench increase. Reaction was mixed on capability: some early users found it solid but not clearly ahead of Kimi/Flash on all tasks (Yuchen Jin’s roundup, scaling01, teortaxesTex), suggesting DeepSeek’s next gains may depend more on RL environment and agent work than raw scale.
  • Microsoft enters with its own reasoning model: Mustafa Suleyman announced MAI-Thinking-1, Microsoft’s first reasoning model “built from scratch,” now available in Foundry. The initial ask from the team is notably practical—Finbarr Timbers specifically requested feedback on tool use—which suggests Microsoft is positioning it as an applied reasoning model rather than just a benchmark entrant.
  • Solar Pro 4 also moved up a tier: Artificial Analysis reported that Upstage’s Solar Pro 4 jumped from 14 to 42 on the Intelligence Index, with especially large gains on agentic and long-context tasks, though still behind the current top frontier and open leaders on both raw score and price.
  • Grok 4.6 在性价比上达到前沿水平:xAI 发布了 Grok 4.6,据称相比 4.5 是重大升级,价格不变。Artificial Analysis 的独立评估将其智能指数评为 61,大致与 GPT-5.6 Sol Max 相当,落后于 Claude Opus/Fable,在代理任务上表现强劲,包括 Terminal-Bench v2.1 上 88.4% 的得分、1753 的 GDPval-AA v2 Elo,以及以远低于竞争对手的成本取得有竞争力的 AA-Briefcase 表现(AA-Briefcase 备注)。Code Arena 的早期竞技场数据也将其在 Web 开发任务上排在 GPT-5.6 Sol 和 Claude Fable 附近。定价是核心主题:AA 强调每 100 万输入/输出 token 的价格为 2/6 美元,明显低于前沿同行,而从业者立即将其视为编码和找 bug 任务的新默认选择(Pawel Huryn,Cognition 在 Devin 中的可用性)。xAI 表示,这些进步来自更长的补充训练运行、重新生成的 SFT 轨迹,以及在编码、网页、CAD 和内核优化上的代理强化学习;他们还报告在长任务期间有更多的自测行为(@kimmonismus 总结)。Elon 还表示 Grok 4.7 已经在开发中,初始训练已完成,计划使用 SpaceX 内部数据进行补充训练。
  • Qwen3.8-Max 开放权重发布:阿里巴巴的 Qwen3.8-Max 以开放权重形式发布,总参数 2.4T,激活参数 95B 的 MoE 模型。社区评论强调其规模、首日服务以及长上下文/代理导向:Yuchen Jin 称其为迄今为止最大的开放权重发布之一;vLLM 提供了首日支持,并为 NVIDIA B300 和 AMD MI355X 提供了特定供应商的 4 位检查点;Together AI 和 Baseten 也宣布立即支持。用户的一个重要注意事项:发布的开放权重版本似乎仅支持文本,初始版本中没有视觉输入(skalskip92)。
  • DeepSeek V4 Pro GA 以低价冲击市场:DeepSeek 的 V4 Pro GA 发布立即引起关注,与其说是“每项基准都最佳”,不如说是经济性。多位观察者强调其定价约为每百万输入 token 0.435 美元,每百万输出 token 0.87 美元(kimmonismus),Cline 称其比 Fable 5 便宜约 57 倍,同时报告相比预览版有显著改进,包括 Terminal Bench 提升 15.8%。对能力的反应不一:一些早期用户发现它表现扎实,但在所有任务上并未明显领先于 Kimi/Flash(Yuchen Jin 的综述、scaling01、teortaxesTex),这表明 DeepSeek 的下一步进展可能更多依赖于 RL 环境和代理工作,而非原始规模。
  • 微软携其推理模型入场:Mustafa Suleyman 宣布了 MAI-Thinking-1,这是微软首个“从零构建”的推理模型,现已在 Foundry 中提供。团队最初的诉求非常务实——Finbarr Timbers 特别要求对工具使用提供反馈——这表明微软将其定位为应用型推理模型,而不仅仅是基准测试的参与者。
  • Solar Pro 4 也上升了一个层级:Artificial Analysis 报告称,Upstage 的 Solar Pro 4 在智能指数上从 14 跃升至 42,在智能体和长上下文任务上尤其取得了巨大进步,但在原始分数和价格方面仍落后于当前顶级前沿模型和开放模型领导者。

Open-Weight Multimodal and Edge Models: Video, Vision, Voice, and Local Inference

开放权重多模态与边缘模型:视频、视觉、语音与本地推理

  • LTX-2.5 and the open video stack keep improving: @RisingSayak highlighted that Lightricks’ LTX-2.5 landed in Diffusers with several practical features that matter for local workflows: joint video + 48 kHz audio generation, prompt-controlled clip length, a 2-pass quality mode, tile rendering for lower memory usage, and preprocessing that re-compresses input images to better match training. Ostris AI Toolkit added support the same day. More broadly, several accounts framed the week as an unusually strong run for open multimedia releases, including MiniMax H3, LTX-2.5, LFM2.5-VL-3B, and North Micro Vision (victormustar, multimodalart).
  • Small VLMs and local multimodal are getting serious: Cohere launched North Micro Vision, an Apache-2.0 open-source small VLM aimed at document understanding, with claims of outperforming Gemma 4 E2B and Ministral 3 3B on a broad visual benchmark mix (results thread). Liquid AI’s LFM2.5-VL-3B was also repeatedly cited as a strong compact vision model, and users demonstrated hybrid local/remote agent stacks—for example, Hermes Agent using DeepSeek V4 Flash for planning plus LFM2.5-VL-3B for local vision.
  • Speech and sign-language releases were unusually substantive: Google DeepMind announced SL2T, a sign-language-to-text system powering ASL input on Android/Pixel 11. The follow-up notes are technically interesting: body pose tracking happens on-device, translation runs server-side, and the system is optimized for real-world constraints like one-handed signing (detail). Separately, Deepgram launched Flux TTS, a low-latency conversational TTS model claiming ~80 ms response time and mid-call adaptation for voice agents.
  • LTX-2.5 和开放视频栈持续改进:@RisingSayak 强调,Lightricks 的 LTX-2.5 已登陆 Diffusers,并具备多项对本地工作流实用的功能:联合视频 + 48 kHz 音频生成、提示词控制的片段长度、双通道质量模式、降低内存使用的分块渲染,以及将输入图像重新压缩以更好匹配训练数据的预处理。Ostris AI Toolkit 也在同一天添加了支持。更广泛地说,多个账号将本周描述为开放多媒体发布的异常强劲的一周,包括 MiniMax H3、LTX-2.5、LFM2.5-VL-3B 和 North Micro Vision(victormustar、multimodalart)。
  • 小型视觉语言模型和本地多模态正在变得严肃:Cohere 发布了 North Micro Vision,一个 Apache-2.0 开源小型视觉语言模型,面向文档理解,声称在广泛的视觉基准组合上优于 Gemma 4 E2B 和 Ministral 3 3B(结果线程)。Liquid AI 的 LFM2.5-VL-3B 也被多次引用为强大的紧凑型视觉模型,用户展示了混合本地/远程智能体栈——例如,Hermes Agent 使用 DeepSeek V4 Flash 进行规划,加上 LFM2.5-VL-3B 进行本地视觉。
  • 语音和手语发布异常充实:Google DeepMind 宣布了 SL2T,一个手语到文本的系统,为 Android/Pixel 11 上的 ASL 输入提供支持。后续说明在技术上很有趣:身体姿态追踪在设备上进行,翻译在服务器端运行,系统针对单手手语等现实世界约束进行了优化(详情)。另外,Deepgram 发布了 Flux TTS,一个低延迟对话式 TTS 模型,声称响应时间约 80 毫秒,并支持语音智能体的通话中适应。

Inference, Compression, and Systems: vLLM, Quantization, CUDA Scheduling, and Ranking Infra

推理、压缩与系统:vLLM、量化、CUDA 调度与排序基础设施

  • vLLM added important infra for giant models and long prompts: vLLM now supports Azure Blob paths for both model loading and KV connectors. The Microsoft/NVIDIA recipe matters operationally: faster weight loading via Dynamo ModelExpress (up to 7.3× faster on H100/A100) and blob-backed KV caching via LMCache + NIXL, trading recomputation for fetches on long-prompt workloads (follow-up).
  • vLLM 为巨型模型和长提示词增加了重要基础设施:vLLM 现在支持 Azure Blob 路径用于模型加载和 KV 连接器。微软/NVIDIA 的配方在操作上很重要:通过 Dynamo ModelExpress 实现更快的权重加载(在 H100/A100 上最高提升 7.3 倍),以及通过 LMCache + NIXL 实现基于 Blob 的 KV 缓存,在长提示词工作负载中用重新计算换取获取(后续跟进)。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近