跳到主内容
精选88elvis论文研究

Meta提出AI研究偏好模型,优化科研Agent实验筛选

Super interesting paper from Meta.

原文
推荐理由

科研Agent的资源调度是核心瓶颈,这篇论文给出了具体的偏好模型架构与量化收益,对做自动化科学探索的团队有直接参考价值。

Super interesting paper from Meta.

来自 Meta 的一篇非常有趣的论文。

Long-horizon research agents are coming.

长周期研究代理即将到来。

But one common problem with research agents today is the lack of originality and how to decide what experiments are worth exploring.

然而,当前研究代理面临的一个常见问题是缺乏原创性,以及如何确定哪些实验值得探索。

An AI research agent can propose far more experiments than it can afford to run, so the problem is not idea generation, it's deciding which candidates get GPU time.

AI 研究代理可以提出远超其运行能力的实验数量,因此问题不在于想法生成,而在于决定哪些候选方案能获得 GPU 时间。

AI Research Preference Models is trained to predict which candidate solution is most promising before any of them execute.

AI 研究偏好模型(AI Research Preference Models)经过训练,可在任何候选方案执行之前预测哪个候选解决方案最具前景。

Two variants, both built on frozen pretrained LLMs. An inference-only model reasons over candidate plans, code, and previously executed solutions. An agentic model additionally runs small-scale pilot experiments before committing budget.

两个变体均基于冻结的预训练大语言模型构建。推理型模型对候选计划、代码和先前执行的解决方案进行推理;智能体模型则在投入预算前额外运行小规模试点实验。

Dropped into the AIRA-dojo agent and measured on AIRS-Bench, average normalized score moves from 0.684 to 0.711 and 0.729.

将其集成到 AIRA-dojo 代理中,并在 AIRS-Bench 上进行评估,平均归一化分数从 0.684 提升至 0.711 和 0.729。

Both variants reach the unguided agent's 24-hour performance in roughly 15 hours, using less than two-thirds of its execution budget, and together set new state of the art on two AIRS-Bench tasks.

两种变体均在约 15 小时内达到了无引导代理 24 小时的性能表现,且使用的执行预算不到其三分之二,并共同在两项 AIRS-Bench 任务上刷新了最新技术水平。

Paper: https://arxiv.org/abs/2608.13940

论文:https://arxiv.org/abs/2608.13940

Chat with Paper: https://academy.dair.ai/papers/ai-research-preference-models-2608.13940

与论文对话:https://academy.dair.ai/papers/ai-research-preference-models-2608.13940

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近