跳到主内容
@wquguru
精选85Rohan Paul论文研究

Meta 论文:量化推理模型因过度怀疑而失败

Paper from Meta shows Quantized reasoning models often lose because they keep do…

原文
发到 X

Paper from Meta shows Quantized reasoning models often lose because they keep doubting a correct answer instead of finishing.

Many of them reason well enough, but compression makes them hesitate at the wrong time.

The problem is that post-training quantization, a way to shrink models after training, can make reasoning models cheaper to run but worse at finishing cleanly.

The authors found that strong quantization does not only make models less capable, since in many failures the model already reached the right answer but then second-guessed itself.

Their core idea is that quantization adds noise at uncertain word choices, so the model becomes more likely to pick words like “wait,” “but,” or “alternatively” that reopen the problem.

They tested this across math, coding, and science tasks using 5 reasoning models, several quantization methods, and model sizes from 1.5B to 32B.

The main result is that aggressive quantization raised overthinking failures up to 52%, while a small penalty on 50 hesitation words cut reasoning length by 12% to 23% and often kept or improved accuracy.

Given compressed models are widely used to save memory and cost, very important to know that a very small decoding fix can stop many of them from wasting tokens and losing answers they already had.

Link – arxiv. org/abs/2606.00206

Title: "Quantized Reasoning Models Think They Need to Think Longer, but They Do Not"

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近