跳到主内容
@wquguru
精选75Rohan Paul论文研究

Meta FAIR 用偏好模型为 AI 实验筛选候选,节省 GPU 算力

AI research agents can generate experiment ideas faster than they can afford to…

原文
发到 X

AI research agents can generate experiment ideas faster than they can afford to run them.

AI研究代理生成实验想法的速度远超其实际运行这些实验的能力。

This Meta FAIR paper attacks that bottleneck: before spending GPU time, a Research Preference Model (RPM) predicts which candidate is worth executing.

这篇Meta FAIR论文针对这一瓶颈:在投入GPU时间之前,一个研究偏好模型(RPM)会预测哪个候选方案值得执行。

AIRA-dojo generates 15 candidate changes at each step.

AIRA-dojo在每一步生成15个候选变更。

The RPM sees their code plus prior results, ranks them, and sends only 1 into the expensive full run.

RPM查看它们的代码及先前结果,进行排序,并仅将1个送入昂贵的完整运行中。

The inference-only RPM uses a frozen LLM.

仅推理的RPM使用冻结的LLM。

The agentic RPM can first spend a 5-minute budget on pilot experiments for extra evidence.

代理型RPM可先花费5分钟的预算进行试点实验以获取额外证据。

Across 20 AIRS-Bench tasks, average normalized score rose from 0.684 with random selection to 0.711 with inference-only RPM and 0.729 with the agentic RPM.

在20个AIRS-Bench任务中,平均标准化分数从随机选择的0.684提升至仅推理RPM的0.711和代理型RPM的0.729。

Both RPM variants matched the baseline's 24-hour score in roughly 15 hours.

两种RPM变体在大约15小时内达到了基线24小时的分数。

– arxiv. org/abs/2608.13940

– arxiv. org/abs/2608.13940

Title: "AI Research Preference Models"

标题:“AI研究偏好模型”

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近