精选85OpenAI模型发布/更新多源精选 ×3
OpenAI审计SWE-Bench Pro:基准已饱和,不再可靠
We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and…
推荐理由
做AI编程评估的同学注意了,SWE-Bench Pro已经饱和,别再依赖它衡量模型能力。OpenAI的审计报告值得细读,建议转向更有效的评估方法。
We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability.
We find the eval to be saturated at a ~70% noise ceiling, and are retracting our previous recommendation that the research community use it as a leading coding eval. https://openai.com/index/separating-signal-from-noise-coding-evaluations/
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力