跳到主内容
@wquguru
精选80Rohan Paul论文研究

AutoLab:前沿模型在长周期研究任务中表现不佳,坚持比聪明更重要

strong AI agents still struggle with long research work because they often fail…

原文
发到 X

strong AI agents still struggle with long research work because they often fail to keep testing and improving.

This paper shows that today’s strongest research agents win less by brilliance than by refusing to stop testing.

The paper proposes AutoLab, a benchmark with 36 tasks where each agent starts from working but weak code and must make it better within a fixed time limit.

The tasks cover system speedups, puzzles, model development, and CUDA kernel work, so the test is not just about writing code once but about managing a long work session.

The authors tested 17 strong models and found that the best results did not mainly come from the first idea being good, but from the model staying active, testing often, and using feedback well.

The best first idea was not the strongest predictor of success; persistence was.

Claude Opus 4.6 led the benchmark not because it always guessed the right move immediately, but because it kept benchmarking and folding empirical feedback into the next attempt.

Several other frontier models failed in a more revealing way: they either quit early with time left on the clock, or thought so long that they ran out of time before submitting anything useful.

– arxiv. org/abs/2606.05080

Title: "AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?"

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近