跳到主内容
@wquguru
精选85OpenAI行业动态多源精选 ×3

OpenAI审计SWE-Bench Pro:30%任务失效,不再推荐

We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and…

原文
发到 X
推荐理由

做AI编码评估的同学注意了,SWE-Bench Pro已被发现30%任务失效,OpenAI不再推荐。赶紧检查你的评估流程是否需要调整。

We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability.

We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as a leading coding eval. https://openai.com/index/separating-signal-from-noise-coding-evaluations/

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →