跳到主内容
@wquguru
精选75elvis论文研究

单谱系提示优化器 NPO 以更少预算追平 GEPA

Interesting paper on prompt optimization.

原文
发到 X

Interesting paper on prompt optimization.

关于提示优化的有趣论文。

They claim that a single-lineage prompt optimizer just matched GEPA on a smaller rollout budget.

他们声称,一个单谱系提示优化器在较小的回滚预算下就与GEPA匹配。

Prompt optimization has been drifting toward heavier machinery, with candidate pools, reflection trees, and Pareto-based selection.

提示优化一直趋向于更重的机制,包括候选池、反思树和基于帕累托的选择。

NPO keeps one lineage. At each iteration it runs the student on the current prompt, collects rollout traces and rewards, and hands a sliding window of recent iterations to a teacher model that rewrites the prompt.

NPO保持单一谱系。在每次迭代中,它在当前提示上运行学生模型,收集回滚轨迹和奖励,并将最近迭代的滑动窗口交给教师模型重写提示。

There is no candidate population and no search tree.

没有候选群体,也没有搜索树。

On the two instruction-following benchmarks it spends 3,500 and 6,800 rollouts against GEPA's 3,593 and 6,871, and it stays broadly comparable across 22 TextArena games.

在两个指令遵循基准上,它花费了3,500和6,800次回滚,而GEPA为3,593和6,871次,并且在22个TextArena游戏中保持大致相当。

The interaction with teacher strength is what makes this interesting. NPO's advantage grows as the teacher model gets stronger, which suggests optimizer-side search complexity has been compensating for weak teacher reasoning all along.

与教师强度的交互是使其有趣的原因。随着教师模型变强,NPO的优势增长,这表明优化器端的搜索复杂性一直在补偿教师推理的不足。

Paper: https://arxiv.org/abs/2608.27266

论文:https://arxiv.org/abs/2608.27266

Chat with Paper: https://academy.dair.ai/papers/naive-prompt-optimization-rethinking-the-need-for-complex-prompt-search-2608.27266

与论文聊天:https://academy.dair.ai/papers/naive-prompt-optimization-rethinking-the-need-for-complex-prompt-search-2608.27266

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近