跳到主内容
@wquguru
精选75Rohan Paul论文研究

新基准揭示AI长程任务短板:17个前沿模型平均通过率仅6.4%

Another paper that so clearly exposes AI’s long-horizon problem.

原文
发到 X

Another paper that so clearly exposes AI’s long-horizon problem.

另一篇论文清晰地揭示了AI在长期任务上的问题。

Long-Horizon-Terminal-Bench, where even frontier agents struggle badly once tasks stretch across hundreds of steps.

Long-Horizon-Terminal-Bench,即便是最前沿的智能体,在任务跨越数百个步骤时也表现挣扎。

Across 17 frontier models on 46 long terminal tasks, the average pass rate is 6.4%.

在46个长期终端任务上测试17个前沿模型,平均通过率仅为6.4%。

And under strict full-completion grading 10 of the 17 models solve zero tasks. Of the runs that fail, 79% end with the clock expiring while the agent is still working.

在严格的完整完成评分下,17个模型中有10个未能解决任何任务。在失败的运行中,79%是因为时间耗尽而智能体仍在工作中。

Even the best model, at 28.3%, leaves 7 of every 10 tasks unfinished.

即便是表现最好的模型,通过率也只有28.3%,意味着每10个任务中有7个未能完成。

The models could perform plenty of locally reasonable steps but could not reliably convert hundreds of those steps into a finished result.

这些模型能够执行大量局部合理的步骤,但无法可靠地将数百个这样的步骤转化为最终成果。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近