跳到主内容
@wquguru
精选86OpenAI行业动态

OpenAI:需建立 AI 不对齐事件披露标准

How we think about the “wiki incident,” where our agents wrote to several intern…

原文
发到 X
推荐理由

AI 治理与合规是行业核心议题,OpenAI 主动提出建立不对齐事件披露新标准,直接影响后续 Agent 部署的安全规范与监管预期,值得从业者关注。

How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models.

我们如何看待“维基百科事件”,即我们的代理向多个互联网网站发送信息:现在是我们定义何时以及如何共享不对齐事件的标准的的时候了,而不仅仅是共享模型的对齐属性。

Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards. This year, we’ve started to see misalignment cause new types of real-world impact.

历史上,我们将不对齐主要视为一个研究问题,并通过系统卡片等研究出版物进行沟通。今年,我们开始看到不对齐导致新型的现实世界影响。

For the Hugging Face incident, where misalignment led to security impact to us and third parties, we followed a traditional security incident response playbook. We immediately started working with Hugging Face to understand what had happened and also disclosed publicly the very next day. Our investigation continues, and we are continuing to notify parties whom our models impacted in less significant ways.

对于 Hugging Face 事件,其中不对齐导致了我们和第三方的安全影响,我们遵循了传统的安全事件响应手册。我们立即开始与 Hugging Face 合作以了解发生了什么,并在第二天公开披露。我们的调查仍在继续,并且我们继续通知那些受到我们模型较小影响的各方。

Prior to the Hugging Face incident, we saw early signs of agents using the internet in unintended ways, as reported in https://openai.com/index/how-we-monitor-internal-coding-agents-misalignment/, https://deploymentsafety.openai.com/gpt-5-6, and https://openai.com/index/safety-alignment-long-horizon-models/. We considered the wiki incident to be an instance of misalignment similar to the ones we’d shared.

在 Hugging Face 事件之前,我们看到代理以非预期方式使用互联网的早期迹象,如 https://openai.com/index/how-we-monitor-internal-coding-agents-misalignment/、https://deploymentsafety.openai.com/gpt-5-6 和 https://openai.com/index/safety-alignment-long-horizon-models/ 中报道的那样。我们认为维基百科事件是与我们分享过的类似不对齐的一个实例。

Our misalignment disclosure practices need to expand for this new phase of model capabilities. We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don’t look like traditional security incidents but could provide insight into AI behavior and future risks. We’re working on a framework and will share it in upcoming weeks, and in parallel we're working with dozens of government regulatory agencies worldwide on these issues.

我们需要扩展不对齐披露实践以适应这一新的模型能力阶段。我们以及更广泛的 AI 社区尚未有明确的报告标准,用于报告在训练、评估和部署期间出现的不对齐情况,包括那些看起来不像传统安全事件但可能提供有关 AI 行为和未来风险的见解的例子。我们正在制定一个框架,并将在接下来的几周中分享它,同时我们与全球数十个政府监管机构在这些问题上进行合作。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近