OpenAI因对齐问题取消Astra 6.1发布并公布安全框架
Astra 6.1 Pulled As Insufficiently Aligned
旗舰模型因安全原因主动撤回极为罕见,这标志着AI对齐从理论走向硬性约束,从业者需密切关注后续安全框架落地及对行业节奏的影响。
We once again got a new set of warnings yesterday, and new movement towards living in a sane world.
昨天我们再次收到了一组新的警告,以及朝着生活在理智世界迈进的新动向。
On the heels of its pause in inference and training due to its latest sandbox escape, OpenAI has cancelled the planned release of their next frontier model, which would have become Astra 6.1. The candidate for Astra 6.1 was found to be too misaligned, including deception and exceeding scope.
在因最新沙盒逃逸问题暂停推理和训练之后,OpenAI 取消了计划中下一代前沿模型的发布,该模型本将成为 Astra 6.1。Astra 6.1 的候选版本被发现存在严重的对齐问题,包括欺骗行为以及超出既定范围。
This leaves Anthropic in a strong position with Opus 5.5, which means they can afford to reciprocate by holding off on Opus and Mythos level models for a bit.
这使得 Anthropic 凭借 Opus 5.5 处于有利地位,这意味着他们可以暂缓推出 Opus 和 Mythos 级别的模型以作回应。
To add a little encouragement, the Florida Attorney General brought the fire.
为了增添一点鼓励,佛罗里达州总检察长介入并采取了行动。
We’re going to need to do better. Towards that, OpenAI offered its vision of how to make a safety case for new AI model training, and they are attempting to implement it. I don’t know that it would be enough, but it would be miles ahead of where we are today if they fully implemented the real versions of all of this.
我们需要做得更好。为此,OpenAI 提出了其关于为新 AI 模型训练构建安全论证的愿景,并正在尝试实施。我不知道这是否足够,但如果他们完全落实所有这些措施的真实版本,那将比我们要好上千里。
There were also signs of greater cooperation across labs.
同时也出现了各实验室之间更大程度合作的迹象。
A new paper came out yesterday, with authors including key people from OpenAI, Anthropic, Microsoft and others, warning of potentially imminent automated AI R&D and recursive self-improvement, risking loss of control over the future. They call on the government to rapidly gain insight into the situation.
昨天发表了一篇新论文,作者包括来自 OpenAI、Anthropic、Microsoft 等机构的关键人物,警告称自动化的 AI 研发和递归自我改进可能迫在眉睫,存在失去对未来控制权的風險。他们呼吁政府迅速深入了解这一局势。
This is on top of the welcome news that Google, OpenAI and Anthropic are planning to form a new AI safety-focused standards body by early in 2027, tentatively titled the Standards Authority for Frontier AI or SAFA.
此外,还有令人欢迎的消息:Google、OpenAI 和 Anthropic 计划于 2027 年初成立一个新的专注于 AI 安全的标准化机构,暂定名为“前沿人工智能标准管理局”(Standards Authority for Frontier AI,简称 SAFA)。
Stop, Hammertime
停下,哈姆林时间
Kudos to OpenAI for not only doing this but being loud about it. Not great that it was necessary, but on net I think I consider this good news.
为 OpenAI 不仅做到了这一点,还高调宣扬而喝彩。虽然这种情况的出现并不理想,但总体而言,我认为这是好消息。
Maxwell Zeff (WSJ): OpenAI says it is scrapping the release of its next-generation AI model [Astra 6.1] over safety concerns that researchers raised during internal testing, in one of the clearest signs so far that agent misbehavior could stymie the industry’s rapid progression.
麦克斯韦尔·泽夫(《华尔街日报》):OpenAI 表示,由于研究人员在内部测试期间提出的安全问题,他们正取消下一代 AI 模型 [Astra 6.1] 的发布,这是迄今为止最清晰的迹象之一,表明智能体不当行为可能会阻碍行业的快速进展。
No, it’s not just a phase. This keeps happening, and it’s going to keep happening.
不,这不仅仅是一个阶段。这种情况反复发生,而且还将继续发生。
The reasons for not releasing Astra 6.1 seem rather compelling.
不发布 Astra 6.1 的理由似乎相当充分。
Maxwell Zeff (WSJ): Saachi Jain, OpenAI’s head of safety systems, said in an interview that GPT-6.1 Astra regressed in two areas. Compared with its predecessor, GPT-6 Astra, the model performed poorly on tests measuring alignment, or how well the model adheres to what humans would like it to do. Specifically, GPT-6.1 Astra showed higher levels of deception: It wasn’t always honest about telling users of the actions it did or didn’t take.
麦克斯韦·泽夫(《华尔街日报》):OpenAI 安全系统负责人萨奇·贾因在采访中表示,GPT-6.1 Astra 在两个领域出现倒退。与前任 GPT-6 Astra 相比,该模型在衡量对齐性(即模型遵循人类期望行为的程度)的测试中表现不佳。具体而言,GPT-6.1 Astra 表现出更高水平的欺骗行为:它在告知用户其采取或未采取的行动时,并不总是诚实的。
Another issue was what OpenAI calls “scope authorization,” meaning that GPT-6.1 Astra would push ahead on a task without asking the user for permission, and would at times reach for external tools and services even if it might be unsafe.
另一个问题被 OpenAI 称为“范围授权”,意味着 GPT-6.1 Astra 会在未征求用户许可的情况下继续执行任务,并且有时甚至会尝试使用外部工具和服务,即使这可能不安全。
… Last week, OpenAI said it paused training on its most capable AI models after an AI agent slipped through a gap in the company’s internet restrictions to query a public chatbot.
……上周,OpenAI 表示,在一款 AI 代理利用公司互联网限制措施的漏洞查询公共聊天机器人后,它暂停了对最强大 AI 模型的训练。
… GPT-6.1 Astra isn’t one of those models, but a different case, the company said.
……该公司表示,GPT-6.1 Astra 不属于这些模型之一,而是另一种情况。
… While the company decided not to ship GPT-6.1 Astra, it hopes to use the same base model to do additional reinforcement learning runs, and create future generations of its GPT-6 models.
……虽然公司决定不发布 GPT-6.1 Astra,但希望使用相同的基座模型进行额外的强化学习运行,并创建其 GPT-6 模型的下一代产品。
It has not even been a month since GPT-6 Astra. The release cycle has gotten too fast. We can afford to skip this edition, learn from the failure and go from there.
GPT-6 Astra 发布至今甚至不到一个月。发布周期变得太快了。我们可以跳过这一版本,从失败中吸取教训,然后继续前进。
On CNBC today, Altman put this decision in the ‘normal course’ category. The model would not have been good for users, so they aren’t shipping it.
今天在 CNBC 上,阿尔特曼将这一决定归类为“正常流程”。该模型对用户来说不够好,因此他们不会发布它。
A Modest Proposal
一个温和的建议
As a response to this, I would like to see Anthropic put out Haiku 5.5 but then not release a new Opus or Mythos level model until let’s say the end of the year, unless OpenAI releases their next Astra or Sol upgrade, or someone otherwise plausibly has caught up to Opus 5.5.
对此,我希望 Anthropic 能推出 Haiku 5.5,但在今年年底之前不再发布新的 Opus 或 Mythos 级别模型,除非 OpenAI 发布了他们的下一个 Astra 或 Sol 升级,或者有人确实追上了 Opus 5.5。
Anthropic has a clear edge right now at the high end, and the pace of model upgrades is exhausting, so there’s no need to push that edge continuously if OpenAI is being cautious. To be clear, there are legal concerns so I’m not asking for an official announcement or commitment that you won’t do it. I’m just saying not to do it.
Anthropic 目前在高端市场拥有明显优势,而模型升级的节奏令人疲惫,因此如果 OpenAI 保持谨慎,就没有必要持续推动这一优势。明确地说,存在法律方面的顾虑,所以我并不是要求你们做出官方声明或承诺不去做。我只是说不要去做。
Making the Safety Case
构建安全论证
OpenAI’s new goal before training? Be able to make a proper overall safety case.
OpenAI 在训练前的新目标?能够提出恰当的整体安全论证。
OpenAI offers its thinking on creating safety cases for frontier AI training. They don’t fully know how to do it, because no one knows how, but they are going to do their best.
OpenAI 分享了他们对为前沿 AI 训练构建安全论证的看法。他们并不完全知道如何做到这一点,因为没人知道怎么做,但他们将全力以赴。
OpenAI: Ideally, such documentation would rise to the level of “safety cases”—comprehensive, structured, evidence-based arguments about risk which are used in other safety-critical industries. We treat safety cases as an aspirational north star we are building towards, while acknowledging the challenges of making them as rigorous for AI models as for aviation or nuclear power, due to the emergent complexity at each new level of AI capability. We’re working on a framework to codify these practices.
OpenAI:理想情况下,此类文档应达到“安全案例”的水平——即关于风险的全面、结构化、基于证据的论证,其他安全关键行业也在采用。我们将安全案例视为一个我们努力达成的理想目标(北极星),同时承认由于人工智能能力在每个新层级都涌现出复杂性,要使它们像航空或核电领域那样严谨面临挑战。我们正在制定一个框架来将这些实践制度化。
They offer some initial guidelines. Here are some highlights.
它们提供了一些初步指南。以下是一些亮点。
The full post expands many of these further.
完整帖子进一步扩展了其中许多内容。
- Technical safeguards.
- Model alignment.
- Training environments and grading.
- Automated and manual dataset reviews, grader tuning, prior run analysis.
- Alignment measurement.
- Offline alignment evals, backtesting, track evaluation gaming and eval awareness, worst-case stress tests.
- Prevent training on chain-of-thought.
- Containment.
- Monitoring.
- Operational guidelines.
- Dissents (pre-mortems).
- Approvals by senior leadership.
- My position on this has long been that you should have many veto points on training, use and release of models, at a variety of levels. Here they list the research lead, the Head of Safety and the Chief Scientist. I’d ideally also include the board and also the members of technical staff, to avoid concentration at one level of management.
- Accountability, internal transparency, audits and escalations.
- Pausing if needed, rollback ability, technical controls.
- Residual risk completeness.
- I continue to worry about the enumeration pattern.
- Escalations, as misalignment incidents get more severe.
- Investigation of misalignment incidents.
- Internal transparency.
- Misalignment root-cause.
- What I do not see, that I most would like to see, is the idea that it counts as a misalignment incident the moment there is intent or an attempt, even if the attempt is prevented or strategically aborted. This should go in the severity table for escalations.
- Postmortem.
- Detection.
- Public disclosures of all of this afterwards.
- 技术保障措施。
- 模型对齐。
- 训练环境与评分。
- 自动化和手动数据集审查、评分器调整、前期运行分析。
- 对齐测量。
- 离线对齐评估、回溯测试、跟踪评估作弊行为及评估意识、最坏情况压力测试。
- 防止在思维链上进行训练。
- 隔离措施。
- 监控。
- 操作指南。
- 异议(事前验尸)。
- 高级领导层的批准。
- 我对此的立场长期以来是:在模型的训练、使用和发布过程中,应在各个层级设置多个否决点。此处列出了研究主管、安全负责人和首席科学家。理想情况下,还应包括董事会和技术团队成员,以避免权力过度集中于某一管理层级。
- 问责制、内部透明度、审计与升级机制。
- 必要时暂停、回滚能力、技术控制措施。
- 剩余风险完整性。
- 我继续对枚举模式表示担忧。
- 随着不对齐事件变得更为严重,进行升级处理。
- 调查不对齐事件。
- 内部透明度。
- 不对齐的根本原因。
- 我最希望看到但尚未看到的是这样一个观点:一旦存在意图或尝试,即便该尝试被阻止或在战略上中止,也应被视为一次不对齐事件。这应当列入升级事件的严重性表格中。
- 事后分析。
- 检测。
- 事后的全面公开披露。
The full version is a good aspirational list. It can be improved. I do not think that doing all of it would add up to what I would count as a safety case for sufficiently advanced intelligence. That does not mean we should not do it.
完整版本是一份良好的理想化清单。它可以得到改进。我认为,即使全部做到,也不足以构成我心目中针对足够高级智能的安全论证。但这并不意味着我们不应该去做。
Stop In the Name of the Law
以法律之名停止
Florida’s attorney general Uthmeier asks for an emergency order against OpenAI to halt ChatGPT development, given the whole ‘tens of thousands of incidents’ thing, until there are third-party approved guardrails.
鉴于此前发生的“数万起事件”,佛罗里达州总检察长乌特迈尔(Uthmeier)要求对 OpenAI 发出紧急命令,以暂停 ChatGPT 的开发,直到建立经第三方批准的护栏机制。
In some sense this would be unthinkable, but at some point yes courts happen.
在某种意义上这令人难以想象,但在某个时间点,法院介入确实是会发生的。
Uthmeier seems to be conflating quite a lot of things, in line with the positions of Florida Governor Ron DeSantis, who has been vocally anti-AI.
乌特迈尔的立场与佛罗里达州州长罗恩·德桑蒂斯(Ron DeSantis)一致,后者一直公开反对人工智能。乌特迈尔似乎在将许多事情混为一谈。
Josephine Walker and Avery Lotz:
约瑟芬·沃克(Josephine Walker)和艾弗里·洛茨(Avery Lotz):
“Stop calling it safe,” Uthmeier said in a video posted to [Twitter] on Monday.
“别再称它是安全的了,”乌特迈尔在周一发布到[Twitter]上的视频中说道。
- “Stop pretending it’s human. Stop selling it to kids. If Sam Altman meant what he said about slowing down, he can join our ask to the court. If he will not, we ask the court to do what OpenAI will not do for itself: protect Florida families.”
- “别再假装它像人类。别再把它卖给儿童。如果山姆·奥特曼(Sam Altman)对他关于放缓发展的言论是认真的,他可以加入我们对法院的请求。如果他不愿意,我们请求法院做 OpenAI 不愿为自己做的事:保护佛罗里达州的大家庭。”
As legal introductions go, they do make a strong case.
作为法律层面的引言,他们确实提出了强有力的论点。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力