OpenAI审计SWE-Bench Pro:30%任务失效,不再推荐
We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and…
推荐理由
做AI编码评估的同学注意了,SWE-Bench Pro已被发现30%任务失效,OpenAI不再推荐。赶紧检查你的评估流程是否需要调整。
We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability.
We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as a leading coding eval. https://openai.com/index/separating-signal-from-noise-coding-evaluations/
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力