自反思循环是否值得?七种方法对比研究给出否定答案
Finally a good paper testing whether self-reflection loops are worth it.
做Agent和推理优化的同学必看,这项研究用严谨的对照实验否定了自反思循环的普遍收益,建议在给Agent加反思步骤前先读一读。
Finally a good paper testing whether self-reflection loops are worth it.
Setup: Seven methods, open models at 1.5B, 3B and 7B, two math benchmarks, 150 questions each.
Every generated token counted, including the ones spent on critiques, reflections, debate turns and checking. Each method compared against repeated sampling at its own measured cost, paired by question, with bootstrap intervals and multiplicity correction.
All 36 comparisons came back with no reliable win for any method. Ten are reliably worse, and every one of those is a method where the model inspects its own output. All 18 self-inspection comparisons are negative.
Self-Refine and a forced Reflexion sit 3.6 to 10.1 points below the baseline at 7B. Reflexion as published never triggered its own retry on the smallest model. It judged itself correct every time and quietly became a single chain of thought.
Worth knowing before you add another critique step to your agent loop.
Paper: https://arxiv.org/abs/2607.28576
Track more trending AI papers in our academy: https://academy.dair.ai/
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力