新基准揭示AI长程任务短板:17个前沿模型平均通过率仅6.4%
Another paper that so clearly exposes AI’s long-horizon problem.
Another paper that so clearly exposes AI’s long-horizon problem.
另一篇论文清晰地揭示了AI在长期任务上的问题。
Long-Horizon-Terminal-Bench, where even frontier agents struggle badly once tasks stretch across hundreds of steps.
Long-Horizon-Terminal-Bench,即便是最前沿的智能体,在任务跨越数百个步骤时也表现挣扎。
Across 17 frontier models on 46 long terminal tasks, the average pass rate is 6.4%.
在46个长期终端任务上测试17个前沿模型,平均通过率仅为6.4%。
And under strict full-completion grading 10 of the 17 models solve zero tasks. Of the runs that fail, 79% end with the clock expiring while the agent is still working.
在严格的完整完成评分下,17个模型中有10个未能解决任何任务。在失败的运行中,79%是因为时间耗尽而智能体仍在工作中。
Even the best model, at 28.3%, leaves 7 of every 10 tasks unfinished.
即便是表现最好的模型,通过率也只有28.3%,意味着每10个任务中有7个未能完成。
The models could perform plenty of locally reasonable steps but could not reliably convert hundreds of those steps into a finished result.
这些模型能够执行大量局部合理的步骤,但无法可靠地将数百个这样的步骤转化为最终成果。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力