跳到主内容
精选85Rohan Paul模型发布/更新

Fable 5 自动化 16.1% 远程工作,AI Agent 能力跃升

CAIS and Scale (AI safety research group) say Fable 5 now automates 16.1% of rea…

原文
推荐理由

做 AI Agent 的同学必看,Fable 5 在真实远程工作基准上翻倍提升,说明 Agent 架构比模型本身更关键,赶紧研究其工具调用和评论家循环设计。

CAIS and Scale (AI safety research group) say Fable 5 now automates 16.1% of real remote-work projects, about 2x Opus 4.8.

Remote Labor Index tests whether an AI can finish paid freelance work well enough for a client to accept it.

Each task comes with briefs, files, and a professional deliverable used as the human baseline.

Fable 5 led the new results, while Opus 4.8 reached 8.3% and GPT-5.5 reached 6.3%.

That 16.1% number still means failure on most tasks, but the direction of improvement is massive. For the context of how large this jump is, the best model scored only 2.5% when RLI launched.

The work is not toy prompting, since tasks include CAD, architecture, animation, audio, data analysis, and web apps. This means the benchmark is testing computer work across messy tools, not narrow text answers.

Fable 5 looked strongest on examples like ring modeling, animation, and bathroom design.

The automated judge ranked models well, but overstated GPT-5.5 by about 2.9x and Opus 4.8 by about 2.3x. The result says AI agents are improving fast, but quality control remains the hard wall.

Another key point is that the model alone is not the whole product anymore. The gains are coming from stronger agent setups: better tool use, full desktop environments, professional software, longer runtimes, and worker-critic loops where 1 agent does the task and another reviews it like a demanding client.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源
V4总结对话能力远超Fable,用户质疑被Anthropic误导
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)原文

相似阅读

另一事件,读法相近