跳到主内容
@wquguru
精选75Rohan Paul论文研究

ASI-Bench:给AI智能体写步骤而非方法名,性能提升显著

Telling a research AI agent which method to use is close to telling it nothing a…

原文
发到 X

Telling a research AI agent which method to use is close to telling it nothing at all.

告诉研究型AI代理使用哪种方法,几乎等同于什么都没告诉它。

What actually carries performance is the procedure, so write the steps rather than the method name.

真正决定性能的是具体步骤,因此要写出步骤,而不是只提方法名称。

ASI-Bench gives agents 60 real research projects across 11 scientific fields. Same goal, same data, same scoring every time. Only the instructions change: full procedure, method name only, or nothing but the objective and the data.

ASI-Bench 为代理提供了跨越11个科学领域的60个真实研究项目。每次的目标、数据和评分标准都相同,唯一变化的是指令:完整步骤、仅方法名称,或只有目标和数据。

Across 18 agent and model combinations, the average score fell from 50.91 with the full procedure to 29.10 with only the method named.

在18种代理和模型组合中,平均得分从提供完整步骤时的50.91分降至仅提供方法名称时的29.10分。

Dropping the method as well cost another 2.5 points. Nearly all the damage comes from losing the steps.

连方法名称也省略,又损失了2.5分。几乎所有的性能损失都源于步骤的缺失。

Method-only prompts were also the most expensive to run, burning 59% more tokens than complete instructions and more than prompts that named no method at all. Naming an approach pins the agent to a direction while still leaving it to rebuild every implementation detail.

仅提供方法名称的提示词运行成本也最高,比完整指令多消耗59%的令牌,甚至比完全不提方法的提示词消耗更多。命名一种方法将代理固定在一个方向上,却仍要它自行重建每个实现细节。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近