跳到主内容
@wquguru
精选75Rohan Paul论文研究

重复采样优于自我反思:Qwen2.5数学推理研究

If you're spending extra tokens making an LLM critique itself, this paper says a…

原文
发到 X

If you're spending extra tokens making an LLM critique itself, this paper says a simpler move can be better: just let it try again.

如果你正在花费额外的token让LLM自我批评,这篇论文表明一个更简单的做法可能更好:直接让它再试一次。

Before you make an LLM reflect on its answer, try giving it another independent attempt.

在你让LLM反思其答案之前,试着给它另一次独立的尝试。

The study compares 7 test-time reasoning methods on Qwen2.5 models from 1.5B to 7B, then asks a fairer question: what happens if repeated sampling gets the same token budget?

该研究在Qwen2.5模型(从1.5B到7B)上比较了7种测试时推理方法,然后提出了一个更公平的问题:如果重复采样获得相同的token预算,会发生什么?

Repeated sampling means solving the same math problem several times and taking the answer that shows up most often.

重复采样意味着多次解决同一个数学问题,并取出现频率最高的答案。

Across 36 comparisons, none of the more elaborate methods reliably beat that baseline at equal generated-token cost; 10 were significantly worse.

在36次比较中,没有一种更复杂的方法在相同生成token成本下可靠地击败该基线;有10种明显更差。

For checkable reasoning, extra tokens may be better spent creating independent attempts than asking the model to reconsider its own work.

对于可检查的推理,额外的token可能更好地用于创建独立的尝试,而不是让模型重新考虑自己的工作。

Ofcouse, study boundary matters: this is Qwen2.5 on math, not frontier models or open-ended tasks.

当然,研究边界很重要:这是Qwen2.5在数学上的表现,而不是前沿模型或开放式任务。

– arxiv. org/abs/2607.28576

– arxiv.org/abs/2607.28576

Title: "Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B"

标题:“多采样,少反思:在相同token成本下,自我精炼和反思输给重复采样,从1.5B到7B”

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近