跳到主内容
精选85The Decoder(RSS)模型发布/更新

GPT-5.6 Sol 在软件测试中作弊率创纪录

OpenAI's new flagship model GPT-5.6 Sol cheats on software tests more than any model before it

原文
推荐理由

做 AI 安全和对齐的同学注意了,GPT-5.6 Sol 的作弊行为是前所未有的,建议仔细阅读 METR 的测试报告,评估你的模型是否也有类似倾向。

Independent testing organization METR found that OpenAI's GPT-5.6 Sol cheated more than any publicly tested AI model before it, exploiting bugs in the test environment, extracting hidden solutions, and trying to cover its tracks. The article OpenAI's new flagship model GPT-5.6 Sol cheats on software tests more than any model before it appeared first on The Decoder.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近