跳到主内容
@wquguru
精选88Rohan Paul论文研究

提示词优化无需复杂搜索树,强教师反馈可替代

Prompt optimization may not need a search tree at all.

原文
发到 X
推荐理由

挑战了提示词优化的主流范式,证明强教师反馈比复杂搜索算法更有效,对 Agent 开发有直接参考价值。

Prompt optimization may not need a search tree at all.

提示词优化可能根本不需要搜索树。

It may depend more on the quality of the teacher and its feedback than on how elaborate the search algorithm is.

它可能更依赖于教师模型的质量及其反馈,而非搜索算法的复杂程度。

This paper tests NPO, which keeps 1 prompt lineage and asks a teacher model to revise it from recent rollout traces and rewards.

本文测试了 NPO,该方法保留一条提示词谱系,并要求教师模型根据最近的 rollout 轨迹和奖励对其进行修订。

Against GEPA, which maintains multiple prompt candidates with Pareto-based selection, NPO reached comparable or better results on IFBench and HotpotQA with slightly fewer rollouts: 3,500 vs 3,593 and 6,800 vs 6,871.

与维持多个提示词候选项并基于帕累托选择进行筛选的 GEPA 相比,NPO 在 IFBench 和 HotpotQA 上取得了相当或更好的结果,且所需的 rollout 次数略少:分别为 3,500 vs 3,593 和 6,800 vs 6,871。

The gap widened as the teacher got stronger.

随着教师模型的增强,差距进一步扩大。

The gap widened as the teacher got stronger.

随着教师模型的增强,差距进一步扩大。

With DeepSeek-V4-Flash and especially GPT-5.5, NPO improved more consistently than GEPA, suggesting that stronger teacher reasoning plus rich feedback can replace some optimizer-side search.

使用 DeepSeek-V4-Flash 尤其是 GPT-5.5 时,NPO 的表现比 GEPA 提升更为一致,这表明更强的教师推理能力加上丰富的反馈可以替代部分优化器侧的搜索。

Those optimized prompts also transferred to other student models, especially within the same model family.

这些经过优化的提示词也迁移到了其他学生模型上,尤其是在同一模型家族内部。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近