LLM评测结果高度依赖评估设置而非模型差异
LLM rankings can be created by evaluation choices as much as model differences,…
这篇论文揭示了当前LLM排行榜的脆弱性,证明评估设置对排名的影响甚至超过模型差异。做模型选型或评测的同学务必参考,避免被单一榜单误导。
LLM rankings can be created by evaluation choices as much as model differences, so one setup should never decide the leaderboard.
大语言模型(LLM)的排名在很大程度上取决于评估方式的选择,而非模型本身的差异,因此单一的设置不应决定排行榜的结果。
The paper keeps the models and 3,679 questions fixed, then changes only ordinary evaluation choices such as prompt format, option order, and scoring method.
该论文保持模型和3,679个问题固定不变,仅改变普通的评估选择,例如提示格式、选项顺序和评分方法。
Those choices move the results a lot.
这些选择会对结果产生显著影响。
gemma4-31b scores anywhere from 31% to 89%, and 4 of the 12 models reach rank 1 under at least 1 valid setup.
gemma4-31b 的得分范围在31%到89%之间,且12个模型中有4个在至少一种有效设置下达到排名第一。
Even more telling, 95.7% of the average gap between neighboring models comes from questions whose answers change when the evaluation setup changes.
更具指示意义的是,相邻模型之间平均差距的95.7%来自于那些在评估设置改变时答案会发生变化的问题。
The biggest source of instability is how answers are scored: generating an answer versus choosing the highest-likelihood option.
不稳定的最大来源在于答案的评分方式:是生成答案还是选择可能性最高的选项。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力