跳到主内容
@wquguru
精选85OpenAI模型发布/更新多源精选 ×3

OpenAI审计SWE-Bench Pro:基准已饱和,不再可靠

We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and…

原文
发到 X
推荐理由

做AI编程评估的同学注意了,SWE-Bench Pro已经饱和,别再依赖它衡量模型能力。OpenAI的审计报告值得细读,建议转向更有效的评估方法。

We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability.

We find the eval to be saturated at a ~70% noise ceiling, and are retracting our previous recommendation that the research community use it as a leading coding eval. https://openai.com/index/separating-signal-from-noise-coding-evaluations/

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近