跳到主内容
精选88Rohan Paul论文研究

LLM评测结果高度依赖评估设置而非模型差异

LLM rankings can be created by evaluation choices as much as model differences,…

原文
推荐理由

这篇论文揭示了当前LLM排行榜的脆弱性,证明评估设置对排名的影响甚至超过模型差异。做模型选型或评测的同学务必参考,避免被单一榜单误导。

LLM rankings can be created by evaluation choices as much as model differences, so one setup should never decide the leaderboard.

大语言模型(LLM)的排名在很大程度上取决于评估方式的选择,而非模型本身的差异,因此单一的设置不应决定排行榜的结果。

The paper keeps the models and 3,679 questions fixed, then changes only ordinary evaluation choices such as prompt format, option order, and scoring method.

该论文保持模型和3,679个问题固定不变,仅改变普通的评估选择,例如提示格式、选项顺序和评分方法。

Those choices move the results a lot.

这些选择会对结果产生显著影响。

gemma4-31b scores anywhere from 31% to 89%, and 4 of the 12 models reach rank 1 under at least 1 valid setup.

gemma4-31b 的得分范围在31%到89%之间,且12个模型中有4个在至少一种有效设置下达到排名第一。

Even more telling, 95.7% of the average gap between neighboring models comes from questions whose answers change when the evaluation setup changes.

更具指示意义的是,相邻模型之间平均差距的95.7%来自于那些在评估设置改变时答案会发生变化的问题。

The biggest source of instability is how answers are scored: generating an answer versus choosing the highest-likelihood option.

不稳定的最大来源在于答案的评分方式:是生成答案还是选择可能性最高的选项。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近