跳到主内容
@wquguru
精选75Rohan Paul论文研究

EdgeBench:最强智能体12小时仅得51.3分,长时程任务仍难突破

Another long-horizon benchmark, another reminder that current AI can work for ho…

原文
发到 X

Another long-horizon benchmark, another reminder that current AI can work for hours and still remain far from mastering the task.

又一个长周期基准测试,又一次提醒我们,当前的人工智能可以连续工作数小时,却仍远未掌握任务。

The strongest agent in EdgeBench scored 51.3/100 after being allowed to work for 12 hours.

在EdgeBench中,最强智能体在允许工作12小时后得分51.3/100。

The benchmark gives agents persistent executable environments and feedback across 134 tasks. They can debug code, run simulations, inspect validation results, revise proofs, and repeatedly submit artifacts to hidden judges.

该基准测试为智能体提供了持久可执行环境和跨134个任务的反馈。它们可以调试代码、运行模拟、检查验证结果、修订证明,并反复向隐藏评判者提交产物。

i.e. they had many of the ingredients we normally say agents need to improve over time.

也就是说,它们具备了我们通常认为智能体随时间改进所需的许多要素。

They still struggle to convert all of that interaction into sustained progress.

但它们仍难以将所有这些交互转化为持续进步。

The study covers roughly 38,000 hours of interaction across 6 task families.

该研究涵盖了6个任务族中约38,000小时的交互。

When performance was averaged across tasks, improvement followed a log-sigmoid curve with R² ≥ 0.997 across all 5 models: slow early progress, a steeper learning phase, then saturation.

当跨任务平均性能时,改进遵循对数S形曲线,所有5个模型的R²≥0.997:早期进展缓慢,随后是更陡峭的学习阶段,最后趋于饱和。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近