跳到主内容
@wquguru
精选75elvis论文研究

Harness Handbook 提升编码智能体规划成功率

Great paper on self-improving agent harnesses.

原文
发到 X

Great paper on self-improving agent harnesses.

(bookmark it)

If you maintain a production agent harness, finding every file behind one behavior is often harder than writing the edit.

Harness Handbook builds a three-level map from runtime behaviors to source locations using static analysis and LLM-assisted structuring.

Its BGPD workflow guides coding agents from the system overview to relevant stages, functions, and files, then verifies every candidate against current source.

Across 60 modification requests on Codex and Terminus-2, handbook guidance raised planning win rates from 28.3% to 38.3% and from 26.7% to 45.6%.

Planner token use fell 12.7% and 8.6%.

File- and symbol-level F1 improved in all 24 comparisons against GPT-5.5 and Opus 4.8 reference plans. Complete localization misses fell by as much as 25.9 points.

This is a strong pattern for coding agents that need to evolve large harnesses without losing scattered or rarely executed behavior.

Paper: https://arxiv.org/abs/2607.13285

Learn to build effective AI agents in our academy: https://academy.dair.ai/

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近