FM-Bench:长周期智能体评估揭示短期基准的局限性
Current agent benchmarks may be ending before the real failures start.
Current agent benchmarks may be ending before the real failures start.
当前的智能体基准测试可能在真正的失败开始之前就结束了。
FM-Bench shows that a model can look strong after 5 years and still finish far behind after 20, so long-running agents need long-running evaluations.
FM-Bench 显示,一个模型在 5 年后可能表现强劲,但在 20 年后仍会远远落后,因此长期运行的智能体需要长期的评估。
FM-Bench has 15 frontier models manage a football club for 20 simulated years, across roughly 340 to 400 decision stops where transfers, contracts, cash, investments, and rival actions keep changing the future.
FM-Bench 让 15 个前沿模型管理一家足球俱乐部长达 20 个模拟年,期间大约有 340 到 400 个决策节点,转会、合同、现金、投资和对手的行动不断改变着未来。
The rankings barely resemble their final shape early on. On seed 1, the year-5 ranking correlated just 0.19 with the final order, and DeepSeek-V4-Pro led at years 5 and 10 but finished 12th.
排名在早期与最终形态几乎没有相似之处。在种子 1 中,第 5 年的排名与最终顺序的相关性仅为 0.19,DeepSeek-V4-Pro 在第 5 年和第 10 年领先,但最终排名第 12。
Competition changes the picture too. In the shared Arena, 10 different models won the league at least once instead of one early leader simply compounding forever.
竞争也改变了局面。在共享竞技场中,有 10 个不同的模型至少一次赢得联赛冠军,而不是由早期的领先者永远累积优势。
What tracked stronger performance was managerial behavior: cutting slow-payoff investments near the end, keeping cash deployed, and renewing contracts earlier. Token use spanned about 7X and still did not order the board.
更能追踪更强表现的是管理行为:在接近结束时削减回报缓慢的投资、保持现金投入运营、以及更早续约。Token 使用量跨度约为 7 倍,但仍未对排行榜进行排序。
So for long-running agents, short task success is a weak proxy for sustained decision quality. The caveat is that the solo board uses 3 seeds and the Arena only 1 shared world.
因此,对于长期运行的智能体来说,短期任务成功只是持续决策质量的弱代理指标。需要注意的是,单人排行榜使用了 3 个种子,而竞技场仅使用 1 个共享世界。
– arxiv. org/abs/2608.18423
– arxiv.org/abs/2608.18423
Title: "FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents"
标题:《FM-Bench:具有竞争性智能体的长周期管理基准》
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力