长周期AI智能体一年任务表现仅为人类27%
This paper is a brutal reality check for long-horizon AI. Give an agent a year o…
This paper is a brutal reality check for long-horizon AI. Give an agent a year of interconnected decisions, delayed feedback, and consequences from its own past actions, and its performance collapses relative to humans.
这篇论文是对长期跨度AI的一次残酷现实检验。给一个智能体一年的相互关联决策、延迟反馈以及自身过去行为带来的后果,其表现相对于人类会大幅下滑。
The researchers tested eight leading models, including GPT-5.6 Sol and Claude Opus 4.8. Yet the best-performing setup, Qwen3.7-Max with Hermes, ended with only 27.3% as much money as the average human participant.
研究人员测试了包括GPT-5.6 Sol和Claude Opus 4.8在内的八款领先模型。然而,表现最佳的配置——Qwen3.7-Max搭配Hermes,最终所获资金仅为普通人类参与者平均水平的27.3%。
A system that finishes a year-long task at barely a quarter of human performance is nowhere near dependable long-horizon execution.
一个完成一年期任务仅达到人类表现四分之一左右的系统,远谈不上可靠的长期执行能力。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力