单谱系提示优化器 NPO 以更少预算追平 GEPA
Interesting paper on prompt optimization.
Interesting paper on prompt optimization.
关于提示优化的有趣论文。
They claim that a single-lineage prompt optimizer just matched GEPA on a smaller rollout budget.
他们声称,一个单谱系提示优化器在较小的回滚预算下就与GEPA匹配。
Prompt optimization has been drifting toward heavier machinery, with candidate pools, reflection trees, and Pareto-based selection.
提示优化一直趋向于更重的机制,包括候选池、反思树和基于帕累托的选择。
NPO keeps one lineage. At each iteration it runs the student on the current prompt, collects rollout traces and rewards, and hands a sliding window of recent iterations to a teacher model that rewrites the prompt.
NPO保持单一谱系。在每次迭代中,它在当前提示上运行学生模型,收集回滚轨迹和奖励,并将最近迭代的滑动窗口交给教师模型重写提示。
There is no candidate population and no search tree.
没有候选群体,也没有搜索树。
On the two instruction-following benchmarks it spends 3,500 and 6,800 rollouts against GEPA's 3,593 and 6,871, and it stays broadly comparable across 22 TextArena games.
在两个指令遵循基准上,它花费了3,500和6,800次回滚,而GEPA为3,593和6,871次,并且在22个TextArena游戏中保持大致相当。
The interaction with teacher strength is what makes this interesting. NPO's advantage grows as the teacher model gets stronger, which suggests optimizer-side search complexity has been compensating for weak teacher reasoning all along.
与教师强度的交互是使其有趣的原因。随着教师模型变强,NPO的优势增长,这表明优化器端的搜索复杂性一直在补偿教师推理的不足。
Paper: https://arxiv.org/abs/2608.27266
论文:https://arxiv.org/abs/2608.27266
Chat with Paper: https://academy.dair.ai/papers/naive-prompt-optimization-rethinking-the-need-for-complex-prompt-search-2608.27266
与论文聊天:https://academy.dair.ai/papers/naive-prompt-optimization-rethinking-the-need-for-complex-prompt-search-2608.27266
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力