IBM新框架BenchDrift:措辞变化致模型分数大幅波动
A model can know the answer and still fail because you asked the same question d…
A model can know the answer and still fail because you asked the same question differently.
模型可能知道答案,但因为提问方式不同而失败。
New IBM paper introduces BenchDrift, an auditing framework for existing benchmarks, that systematically rephrases the same questions without changing their answers, then measures how much a model's score moves purely because of wording.
IBM的新论文介绍了BenchDrift,一个针对现有基准的审计框架,它系统地改写相同问题而不改变答案,然后衡量模型得分仅因措辞变化而波动的幅度。
Across 8 models and 3 benchmarks, the gap between best-case and worst-case accuracy averaged 74.7 percentage points.
在8个模型和3个基准测试中,最佳与最差情况下的准确率差距平均为74.7个百分点。
Stronger models were actually more exposed.
更强的模型实际上暴露得更多。
For every model-benchmark pair above 60% baseline accuracy, rephrasing broke more correct answers than it recovered.
对于每个基线准确率超过60%的模型-基准组合,改写导致更多正确答案被破坏而非恢复。
And confidence did not solve this.
而置信度并未解决这一问题。
Even in the highest-confidence bucket, 18.5% of correct answers were lost after meaning-preserving rewording.
即使在最高置信度区间,经过保持语义的改写后,仍有18.5%的正确答案丢失。
So if your model selection depends on small score differences, test several equivalent phrasings or report the range.
因此,如果你的模型选择依赖于微小的分数差异,应测试多种等价表述或报告分数范围。
– arxiv. org/abs/2608.11694
– arxiv.org/abs/2608.11694
Title: "The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance"
标题:“措辞效应:量化LLM基准性能中的双向漂移”
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力