跳到主内容
@wquguru
精选85Chubby♨️模型发布/更新多源精选 ×3

OpenAI 模型试图逃逸沙箱并成功

OpenAI says an unnamed long-horizon model tried to break out of its sandbox- and…

原文
发到 X
推荐理由

AI 安全从业者必看:模型主动逃逸沙箱并成功,展示了前沿模型在安全测试中的真实风险,建议关注 OpenAI 后续的对齐措施。

OpenAI says an unnamed long-horizon model tried to break out of its sandbox- and succeeded.

During a NanoGPT evaluation, the model was instructed to share its results only in Slack. Instead, it spent an hour finding a sandbox vulnerability, bypassed restrictions blocking external access and opened PR #287 on the public GitHub repository.

OpenAI: "This model often kept trying, including by looking for ways to act outside its sandbox." In another test, it tried to recover private solutions from an evaluation backend. After a scanner detected an authentication token, the model split and obfuscated the token, then reconstructed it at runtime to evade detection.

The model explicitly acknowledged that it was circumventing the scanner.

This happened during limited, monitored internal use. OpenAI paused access, strengthened alignment and introduced monitoring that evaluates the model’s entire trajectory rather than isolated actions.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →