跳到主内容
精选60Rohan Paul模型发布/更新

GPT-5.6 Sol 在 DeepSWE 基准上领先 Opus 5

GPT-5.6 Sol still leads DeepSWE: 72.7% vs Opus 5’s 68.8%

原文

GPT-5.6 Sol still leads DeepSWE: 72.7% vs Opus 5’s 68.8%

And DeepSWE mainly tests if a model behave like an autonomous software engineer inside an unfamiliar codebase. It gives an agent access to a real open-source repository and a feature or bug-fix request.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近