耶鲁等机构论文:前沿模型物理基准测试质量成瓶颈
New Yale Univ + other top lab paper shows frontier models are already close to m…
这篇论文直接挑战了用单一基准分数评估模型能力的惯例,通过严谨的人工审计揭示了评测体系的系统性偏差,对理解模型真实水平极具参考价值。
New Yale Univ + other top lab paper shows frontier models are already close to maxing out today’s closed-ended physics benchmarks.
耶鲁大学及其他顶尖实验室的新论文表明,前沿模型已接近当前封闭式物理基准测试的上限。
And many apparent failures come from bad questions, wrong reference answers, and brittle graders explain most audited physics failures, making benchmark quality the new bottleneck for measuring frontier models.
许多看似失败的情况源于糟糕的问题、错误的答案参考以及脆弱的评分器,这解释了大多数经过审计的物理测试失败案例,使得基准质量成为衡量前沿模型能力的新瓶颈。
The researchers had physicists re-check model failures across 6 popular physics benchmarks instead of trusting the original scores.
研究人员让物理学家重新检查了6个流行物理基准测试中的模型失败案例,而不是盲目信任原始分数。
In 250 rejected cases from 4 audited benchmark subsets, only 12 were actual model mistakes.
在4个经过审计的基准子集中被拒绝的250个案例中,仅有12个是模型的实际错误。
The other 238 came from bad questions, wrong reference answers, or graders rejecting correct answers.
其余238个案例源于糟糕的问题、错误的答案参考或评分器拒绝了正确答案。
After expert review, GPT-5.6-Sol’s measured HLE-Physics score rose from 47.3% to 78.7%.
经过专家审查后,GPT-5.6-Sol 在 HLE-Physics 基准上的得分从 47.3% 上升至 78.7%。
So a low physics benchmark score can badly underestimate what a frontier model can actually solve.
因此,较低的物理基准分数可能会严重低估前沿模型实际能够解决的问题。
But that does not mean these models can reliably do physics research: the authors’ agents still failed to fully solve any of the open theoretical-physics problems they tried.
但这并不意味着这些模型能够可靠地进行物理研究:作者的智能体仍未能完全解决它们尝试的任何开放理论物理问题。
but it does mean that do not treat benchmark scores as clean ground truth anymore for AI's Physics capability.
但这也意味着我们不应再将基准分数视为 AI 物理能力的干净事实标准。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力