OpenAI暂停前沿模型训练以排查智能体越权漏洞
OpenAI halts frontier-model training amid string of agent misalignment incidents
前沿模型因安全对齐问题主动暂停训练,直接反映当前 Agent 能力边界与安全治理现状,值得从业者关注其后续修复方案。
OpenAI says it has paused all internal training of "our most capable models" as it continues what CEO Sam Altman is calling "an extensive and ongoing review related to our agents’ use of internet access during training and evaluation."
OpenAI 表示,它已暂停所有内部训练“我们最强大的模型”,同时继续进行 CEO Sam Altman 所称的“与我们代理在训练和评估期间使用互联网访问相关的广泛且持续的审查。”
The company revealed the pause in a report about a so-called misalignment incident in which an agent attempted to exploit a gap in Internet-access restrictions during a routine research task during training. OpenAI says that improper DNS filtering allowed the agent to attempt to break out of its sandbox and access the wider Internet when asked for biographical details about a blogger.
该公司在一篇关于所谓“对齐偏差”事件的报告中披露了此次暂停,在该事件中,一个代理试图在训练期间的常规研究任务中利用互联网访问限制中的漏洞。OpenAI 称,不当的 DNS 过滤允许该代理在被要求提供某位博主的传记细节时,尝试突破其沙箱环境并访问更广泛的互联网。
OpenAI says the agent was only able to access the company's offline web cache and that it has implemented additional multi-layered blocking controls to prevent similar incidents in the future. Despite that, though, the company says it has decided to "pause all other training, evaluation, and inference with tool-use" for this frontier model "until we have both validated that the gap is resolved and performed additional red-teaming of the system."
OpenAI 表示,该代理仅能访问公司的离线网页缓存,并且已实施额外的多层阻止控制措施,以防止未来发生类似事件。尽管如此,该公司表示已决定“暂停该前沿模型的所有其他使用工具的训练、评估和推理”,直到“我们验证该漏洞已得到解决并对系统进行了额外的红队测试”。
Read full article
阅读全文
Comments
评论
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力