IBM新研究:基准分数受措辞影响,强模型更脆弱
Interesting new research from IBM.
Interesting new research from IBM.
IBM的新研究很有趣。
If you pick models from benchmark deltas, some of that delta belongs to the phrasing rather than the model.
如果你根据基准测试的差异来选择模型,那么部分差异属于措辞而非模型本身。
BenchDrift generates meaning-preserving variations of benchmark problems along linguistic, referential, pragmatic, and structural axes, holding the answer fixed, then measures how often correctness flips.
BenchDrift沿语言、指称、语用和结构轴生成保持意义不变的基准问题变体,同时固定答案,然后测量正确性翻转的频率。
Phrasing sensitivity does not fade as models improve. It changes sign. Weak models gain more from rephrasing than they lose, while strong models lose far more than they gain, so the top models on a benchmark are the ones whose scores depend most on the wording they happened to receive.
措辞敏感性并不会随着模型改进而消失。它会改变符号。弱模型从改写中获得的收益多于损失,而强模型损失远大于收益,因此基准测试中排名靠前的模型,其得分最依赖于它们碰巧收到的措辞。
Fragility also belongs to the rephrasing. Across eight models on GSM8K, MMLU, and MATH-Hard, they largely agree on which rephrasings cost the most correct answers even while differing in how much they drift overall.
脆弱性也属于改写。在GSM8K、MML U和MATH-Hard上对八个模型进行测试,它们大体上一致认为哪些改写会损失最多的正确答案,即使在整体漂移程度上有所不同。
Rephrasing breaks answers models were confident about, whether the problem gets shorter or longer.
改写会破坏模型原本有信心的答案,无论问题变短还是变长。
Paper: https://arxiv.org/abs/2608.11694
论文:https://arxiv.org/abs/2608.11694
Track more trending AI papers in our academy: https://academy.dair.ai/
在我们的学院中追踪更多热门AI论文:https://academy.dair.ai/
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力