Meta FAIR 用偏好模型为 AI 实验筛选候选,节省 GPU 算力
AI research agents can generate experiment ideas faster than they can afford to…
AI research agents can generate experiment ideas faster than they can afford to run them.
AI研究代理生成实验想法的速度远超其实际运行这些实验的能力。
This Meta FAIR paper attacks that bottleneck: before spending GPU time, a Research Preference Model (RPM) predicts which candidate is worth executing.
这篇Meta FAIR论文针对这一瓶颈:在投入GPU时间之前,一个研究偏好模型(RPM)会预测哪个候选方案值得执行。
AIRA-dojo generates 15 candidate changes at each step.
AIRA-dojo在每一步生成15个候选变更。
The RPM sees their code plus prior results, ranks them, and sends only 1 into the expensive full run.
RPM查看它们的代码及先前结果,进行排序,并仅将1个送入昂贵的完整运行中。
The inference-only RPM uses a frozen LLM.
仅推理的RPM使用冻结的LLM。
The agentic RPM can first spend a 5-minute budget on pilot experiments for extra evidence.
代理型RPM可先花费5分钟的预算进行试点实验以获取额外证据。
Across 20 AIRS-Bench tasks, average normalized score rose from 0.684 with random selection to 0.711 with inference-only RPM and 0.729 with the agentic RPM.
在20个AIRS-Bench任务中,平均标准化分数从随机选择的0.684提升至仅推理RPM的0.711和代理型RPM的0.729。
Both RPM variants matched the baseline's 24-hour score in roughly 15 hours.
两种RPM变体在大约15小时内达到了基线24小时的分数。
– arxiv. org/abs/2608.13940
– arxiv. org/abs/2608.13940
Title: "AI Research Preference Models"
标题:“AI研究偏好模型”
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力