跳到主内容
@wquguru
精选86Import AI(RSS)论文研究

METR研究AI加速差异及SPADE、Hawkeye技术解析

Import AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye

原文
发到 X
推荐理由

涵盖METR关于AI加速效应的实证研究及两项前沿工程方法(SPADE与Hawkeye),对理解当前AI能力边界与训练范式有重要参考价值。

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe.

欢迎来到 Import AI,这是一份关于人工智能研究的通讯。Import AI 的运行依赖于 arXiv、卡布奇诺咖啡以及读者的反馈。如果您希望支持我们,请订阅。

Subscribe now

立即订阅

AI is accelerating some types of progress but not others:

人工智能正在加速某些类型的进步,但并非所有类型:

…A nice METR study lays out where acceleration is showing up…

……METR 的一项研究清晰地展示了加速效应出现在哪些领域……

Here’s a little analysis from METR which looks at where AI may be accelerating different types of science and technology. The study looks at three different areas: cyber, math, and AI research, and finds that AI has contributed a lot to cyber, a little bit to math, and it’s hard to say for AI.

以下是 METR 的一份简要分析,探讨了人工智能可能在哪些方面加速不同类型的科学与技术发展。该研究考察了三个不同领域:网络安全、数学和人工智能研究,发现人工智能对网络安全贡献巨大,对数学有一定贡献,而对人工智能本身的影响则难以定论。

Where have LLMs actually made a difference to scientific discovery?

大型语言模型(LLMs)究竟在科学发现的哪些方面产生了实际影响?

  • Cyber vulnerabilities: Major acceleration. “The rate of vulnerabilities reported across many projects has dramatically accelerated in 2026 compared with 2025, both for specific projects (cURL, OpenSSL, Firefox, and Microsoft) and for aggregate vulnerability databases (the US NVD, and OSV)”.
  • Mathematics research: Minor acceleration, but harder to measure. “AI is clearly contributing to more work being done (arXiv submissions have doubled in some areas in less than 12 months) but quantifying the value of those contributions is difficult.” Some math problems from prestigious lists have been solved, e.g., “the Jacobian conjecture from Smale’s list, Problem 44 from Green’s list (the halving sieve), and the sofic half of Green’s Problem 100”. However, it may be too early to determine how sustained a trend this is.
  • Optimization of AI research: No measurable acceleration. When you look at algorithmic progress across seven significant problem areas (CIFAR-10, Hutter compression, Gurobi mixed-integer programming, MIPLIB, nanoGPT, Stockfish, and the matrix-multiplication exponent) there are a couple of these where LLM-attributable contributions have happened (nanoGPT, CIFAR-10), though the rate of increase of usage of AI here is a lot less than with cybersecurity and mathematics.
  • 网络安全漏洞:显著加速。“与 2025 年相比,2026 年许多项目中报告的漏洞率大幅加速,这既适用于特定项目(cURL、OpenSSL、Firefox 和微软),也适用于综合漏洞数据库(美国国家漏洞数据库 NVD 和 OSV)。”
  • 数学研究:轻微加速,但更难衡量。“人工智能显然促成了更多工作的完成(在某些领域,arXiv 的投稿量在不到 12 个月内翻了一番),但量化这些贡献的价值十分困难。”一些来自著名榜单的数学问题已被解决,例如“Smale 列表中的雅可比猜想、Green 列表中的第 44 题(减半筛法),以及 Green 第 100 题中所谓的‘sofic’部分”。然而,现在判断这一趋势是否具有持续性可能还为时过早。
  • 人工智能研究的优化:无可测量的加速。当审视七个重要问题领域的算法进展时(CIFAR-10、Hutter 压缩、Gurobi 混合整数规划、MIPLIB、nanoGPT、Stockfish 和矩阵乘法指数),其中少数几个领域出现了可归因于大型语言模型的贡献(nanoGPT、CIFAR-10),尽管此处人工智能使用量的增长率远低于网络安全和数学领域。

Why this matters – differential acceleration: This paper highlights how AI is causing advances in some parts of science and technology, but the effect isn’t unified across fields, rather there are pockets of lumpy acceleration (e.g., cyber) and areas where progress is more gradual (math, AI). My suspicion is that acceleration happens when models go through some kind of ineffable phase change for a given skill, as has evidently happened with day-to-day coding (2025), and cyber (2026). The key question is whether we are going to see phase changes in other parts of science and technology or if we won’t.

为何重要——差异化加速:本文强调了人工智能如何推动科学和技术某些领域的进步,但这种影响并非在所有领域统一发生,而是存在加速的“斑块”(例如网络安全),以及进展更为渐进的领域(数学、人工智能)。我的推测是,当模型在特定技能上经历某种难以言喻的相变时,加速就会发生,这在日常编程(2025年)和网络安全(2026年)中显然已经出现。关键问题在于,我们是否会在科学和技术的其他部分看到相变,还是不会。

Read more: Research note: Have We Seen an Acceleration in Discoveries? (METR).

阅读更多:研究笔记:我们发现发现加速了吗?(METR)。

Automating environment generation with SPADE:

使用 SPADE 自动化环境生成:

…A crude form of RSI bootstrapping via increasing data breadth…

……一种通过增加数据广度进行的 RSI 自举粗糙形式……

A multi-university group of researchers have built SPADE, Self-Play in Adaptive Synthetic Executable Environments. SPADE is a “general framework for co-evolving environments synthesis and agentic capability through self-play”, and works as a way to generate synthetic data in the form of game-like environments which LLMs can subsequently be trained in, allowing developers to use a powerful model to bootstrap the creation of data that can then be used to further refine that same model.

一个多大学的研究团队构建了 SPADE(自适应合成可执行环境中的自我博弈)。SPADE 是一个“通过自我博弈共同进化环境合成与智能体能力的通用框架”,其工作方式是通过生成类似游戏环境的合成数据,供大型语言模型随后进行训练,使开发人员能够利用强大的模型来引导数据的创建,而这些数据又可用于进一步精炼该模型本身。

Who did it: SPADE was developed by researchers with the University of Washington, Stanford University, Northeastern University, Carnegie Mellon University, Massachusetts Institute of Technology, National University of Singapore, Seoul National University, Stevens Institute of Technology, and the University of Chicago.

谁做的:SPADE 由华盛顿大学、斯坦福大学、东北大学、卡内基梅隆大学、麻省理工学院、新加坡国立大学、首尔国立大学、史蒂文斯理工学院和芝加哥大学的研究人员开发。

How it works: SPADE has an LLM alternate between generating executable training environments (e.g., puzzles where a system needs to solve a simulated genetic problem in a biology lab) and having an LLM try to solve them. SPADE has two key roles for the model being used:

工作原理:SPADE 让大型语言模型交替执行两项任务:生成可执行的训练环境(例如,需要系统在生物实验室中解决模拟遗传问题的谜题),并尝试让大型语言模型解决这些环境。SPADE 对所使用的模型有两个关键角色设定:

  • Environment Designer; writes complete, long-horizon training environments as executable code.
  • Reasoning Agent; learns to act in the environments. The reward for the reasoning agent is estimated using the gap between its reward with and without privileged hints. A privileged hint (h) is “task-relevant information that the Environment Designer attaches to an environment (for example, a partial solution sketch or a key structural observation); revealing h to the Reasoning Agent makes the environment easier to solve, and the gap in Reasoning Agent return with versus without h defines the Environment Designer’s hint-based regret reward”.
  • 环境设计师;编写完整的、长周期的训练环境作为可执行代码。
  • 推理智能体(Reasoning Agent);学习在环境中采取行动。推理智能体的奖励是通过其带有特权提示和不带特权提示时的奖励之差来估计的。特权提示(h)是“环境设计师附加到环境中的与任务相关的信息(例如,部分解决方案草图或关键结构观察结果);向推理智能体揭示 h 会使环境更容易解决,而推理智能体在有 h 和没有 h 情况下的回报差距定义了环境设计师基于提示的遗憾奖励。”

It works at the 30B scale: The authors train three Qwen3 backbones to test out SPADE: Qwen3-4B-Instruct-2507, Qwen3-8B, and Qwen3-30B-A3B-Instruct-2507. Unsurprisingly, Qwen3-30B works the best. Each model is tuned via GRPO for 400 rollouts of 25 environments each, then assessed against a variety of benchmarks including AIME, GPQA, LCB, and environments within Reasoning Gym. They generate two types of environments – game environments, and tool-use environments. SPADE improves performance on both.

它在 30B 规模上运行:作者训练了三个 Qwen3 骨干网络以测试 SPADE:Qwen3-4B-Instruct-2507、Qwen3-8B 和 Qwen3-30B-A3B-Instruct-2507。不出所料,Qwen3-30B 表现最佳。每个模型均通过 GRPO 进行微调,共 400 次 rollout,每次包含 25 个环境,随后在包括 AIME、GPQA、LCB 以及 Reasoning Gym 内的各种基准和环境进行评估。他们生成了两种类型的环境——游戏环境和工具使用环境。SPADE 在这两方面都提升了性能。

For games, “at 30B-A3B, SPADE reaches a suite average of 58.3: +8.1 over base and +5.3 over the strongest fixed-environment baseline”. For tools, they see the same significant boost: “the same recipe applied to tool-use environment design improves every backbone”.

对于游戏,“在 30B-A3B 规模下,SPADE 达到套件平均分为 58.3:比基线高出 8.1,比最强的固定环境基线高出 5.3”。对于工具,他们也看到了同样的显著提升:“将相同的配方应用于工具使用环境设计,提升了所有骨干网络的性能。”

Why this matters – part of RSI: This is basically a form of fancy synthetic data generation, letting researchers use whatever powerful model they have to hand to generate a more diverse set of training environments for another model to train against. I suspect that you could repeatedly swap out the powerful model (e.g., toggling between different frontier models from different companies) to increase the diversity of your environment generation. This kind of technique makes it a lot cheaper to build big, broad datasets to use to train models on. Though, as the authors note, it doesn’t allow models to bootstrap themselves massively beyond the imaginative capabilities of the base model used for environment generation.

为何重要——RSI 的一部分:这基本上是一种高级合成数据生成形式,允许研究人员利用手头任何强大的模型为另一个模型生成更多样化的训练环境集。我怀疑你可以反复替换掉那个强大的模型(例如,在不同公司的不同前沿模型之间切换),以增加环境生成的多样性。这种技术使得构建大型、广泛的数据集以用于模型训练变得更加便宜。不过,正如作者所指出的,它并不能让模型超越用于环境生成的基础模型的想象力能力,从而实现大规模的自我引导提升。

“By representing environments as Python programs with a Gym-style interface, the framework unifies single-turn reasoning and multi-turn agentic tasks, and turns environment design into a learnable, RL-trained component of post-training, enabling continual open-ended self-improvement,” they write.

他们写道:“通过将环境表示为具有 Gym 风格接口的 Python 程序,该框架统一了单轮推理和多轮智能体任务,并将环境设计转变为后训练中一个可学习、经强化学习训练的组件,从而实现持续的开放式自我改进。”

Read more: SPADE: Self-Play in Adaptive Synthetic Executable Environments (arXiv).

阅读更多:SPADE: Self-Play in Adaptive Synthetic Executable Environments (arXiv)。

Get the code here, including model checkpoints: SPADE (spade-rl, GitHub).

在此获取代码,包括模型检查点:SPADE (spade-rl, GitHub)。

Building better GPU kernels with Hawkeye:

使用 Hawkeye 构建更优秀的 GPU 内核:

…Well-documented unit tests can boost performance of kernel-writing agents…

……完善的单元测试可以提升编写内核的 AI 代理的性能……

Researchers with Harvard, Stanford, Together AI, and Caltech have built Hawkeye, software to make it easier for agents to learn how to write well-optimized kernels for specific types of GPU hardware. Systems like Hawkeye are important because they’re essentially tools that AI systems can use to boost their performance on tasks related to AI R&D, like optimizing the performance of a given AI system on a given piece of hardware. The goal of the project is to answer the question “how can we make coding agents hardware-aware with minimal expert intervention?”

来自哈佛大学、斯坦福大学、Together AI 和加州理工学院的研究人员构建了 Hawkeye,这是一款旨在帮助 AI 代理学习如何为特定类型的 GPU 硬件编写高度优化内核的软件。像 Hawkeye 这样的系统之所以重要,是因为它们本质上是 AI 系统可用于提升其在 AI 研发相关任务(如在给定硬件上优化给定 AI 系统的性能)表现的工具。该项目的目标是回答这样一个问题:“我们如何在最少专家干预的情况下,使编码代理具备硬件感知能力?”

Hawkeye is “an open-source framework that grounds autonomous kernel generation in a minimal and comprehensive taxonomy”, the researchers write. It “demonstrates that minimally supervised coding agents can exploit architecture-specific hardware features and reduce the overhead of supporting emerging hardware accelerators”.

研究人员写道,Hawkeye 是“一个开源框架,它将自主内核生成建立在最小且全面的分类体系之上”。它“证明了在极少监督下的编码代理能够利用特定架构的硬件特性,并降低对新兴硬件加速器的支持开销”。”

The key contribution of Hawkeye is that it “introduces a generalizable, minimal, and comprehensive taxonomy of unit tests that enables coding agents to scale test-time compute more effectively and generate hardware-aware kernels”. In other words, it basically ships as a well-curated set of information about different hardware platforms and the optimization strategies to use on them, packaged up as unit tests.

Hawkeye 的关键贡献在于它“引入了一种可泛化、最小且全面的单元测试分类体系,使编码代理能够更有效地扩展测试时的计算资源,并生成具备硬件感知能力的内核”。换句话说,它基本上打包了一套精心策划的信息,涵盖不同硬件平台及其适用的优化策略,并以单元测试的形式呈现。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件