跳到主内容
@wquguru
精选70Rohan Paul论文研究

长时程智能体可靠性未随模型提升而到来

Long-horizon agent reliability has not arrived yet with better models.

原文
发到 X

Long-horizon agent reliability has not arrived yet with better models.

长时程智能体的可靠性并未随模型改进而实现。

On WeaveBench's 114 hybrid GUI-CLI tasks, the best officially reported pass rate is 41.2%.

在WeaveBench的114项混合GUI-CLI任务中,官方报告的最佳通过率为41.2%。

Current models can handle local steps, but need explicit audited state to keep the full task on track.

当前模型能处理局部步骤,但需明确审计的状态来保持整个任务不偏离轨道。

Long-horizon agents can solve individual steps and still fail because history becomes unreliable about what is finished, what failed, and what remains.

长时程智能体能解决单个步骤,但仍可能失败,因为历史记录对已完成、失败及剩余任务变得不可靠。

LongHorizon-Harness treats this as a task-state problem.

LongHorizon-Harness将此视为任务状态问题。

this paper argues that reliability over that span depends as much on the harness around the model as on the model itself.

本文认为,在该跨度上的可靠性既取决于模型本身,也取决于围绕模型的框架。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近