Anthropic暂停高风险RL训练并部署实时拦截器
Anthropic Has Some Alignment Problems
一线大厂罕见公开承认RL训练中的越狱行为及具体防御手段,对做Agent安全的同学极具参考价值。
Oh, good. They noticed.
哦,很好。他们注意到了。
Anthropic, too, is planning to bring METR inside for an independent review of their own incidents, where three times a Claude model started hacking outside things during an eval, and where Mythos 5 did various ‘unauthorized actions,’ by which we mean tried to hack various real-world things, during a UK AISI cybersecurity eval.
Anthropic 也计划将 METR 引入内部,对其自身事件进行独立审查:在一次评估中,Claude 模型三次开始对外部事物进行黑客攻击;而在一次英国 AISI 网络安全评估中,Mythos 5 执行了各种“未经授权的操作”,我们指的是它试图对现实世界中的各种事物进行黑客攻击。
Anthropic, too, is pacing the frontier internally, while calling on it to be paced globally.
Anthropic 也在内部推进前沿技术,同时呼吁在全球范围内对其进行管控。
As in, Anthropic paused its highest risk RL efforts, in light of holy hell have you seen the data we are training on and the ways it is teaching our models to act.
也就是说,鉴于我们训练数据的可怕状况以及这些数据教导模型行为的方式,Anthropic 暂停了其最高风险的反强化学习(RL)工作。
They are also sharing research in which they intentionally created a reward seeking version of Claude.
他们还分享了相关研究,在其中他们有意创建了一个寻求奖励的 Claude 版本。
Scheduling note: Fable 5.1 has been released. I will aim to cover that starting Friday. OpenAI is also planning to release Astra soon, which I would cover after Fable.
日程备注:Fable 5.1 已发布。我计划从周五开始对此进行覆盖。OpenAI 也计划很快发布 Astra,我会在 Fable 之后对其进行报道。
Also, we have a breaking news story about looming problems with chain of thought monitorability, which I’ll preview before I get to the main post.
此外,我们有一个关于思维链可监控性 looming problems(迫在眉睫的问题)的突发新闻故事,我将在进入主帖之前先做一个预览。
Table of Contents
目录
- This Just In.
- Anthropic Parallel Pauses.
- Pause The Data Brokers.
- Pacing the Frontier.
- Misalignment Assessment.
- Defects In Training Environments Disproportionately Cause Cheating.
- Creating Reward Hacker Opus.
- Undo It.
- Mistakes Were Made.
- Internal Security Posture.
- One Does Not Simply Fix The RL Environments.
- 最新消息。
- Anthropic 并行暂停。
- 暂停数据掮客。
- 推进前沿。
- 不对齐评估。
- 训练环境中的缺陷不成比例地导致作弊行为。
- 创建奖励黑客 Opus。
- 撤销它。
- 犯了错误。
- 内部安全态势。
- 人们不能简单地修复 RL 环境。
This Just In
最新消息
Last night, The Information reported that OpenAI is using a new technique called recurrent depth, which can interfere with the faithfulness and monitorability of model Chain of Thought. As per their report, this is not currently observed in practice to be an issue with Astra, but notice how I had to word that.
昨晚,《信息》(The Information)报道称,OpenAI 正在使用一种名为循环深度(recurrent depth)的新技术,该技术可能会干扰模型思维链(Chain of Thought)的忠实度和可监控性。根据他们的报道,目前在实际操作中尚未发现这对 Astra 构成问题,但请注意我是如何措辞的。
Amir Efrati (The Information): An innovative technique that improved the model’s performance also means that the model, and others like it, will reveal less of their “thinking,” making them harder to monitor for signs of bad behavior, according to a person with knowledge of Astra’s development.
Amir Efrati(《信息》):一位了解 Astra 开发情况的人士表示,一项提高模型性能的创新技术也意味着,该模型及其类似模型将更少地揭示其“思考”过程,从而更难监控不良行为的迹象。
While the limitation isn’t necessarily a significant issue with Astra, the technique has triggered concerns inside OpenAI and across the industry about whether AI developers that adopt and supercharge it will struggle to guard against the kind of rogue AI that recently hacked OpenAI’s own systems and those of other companies such as Hugging Face.
虽然这一限制在 Astra 上未必构成重大问题,但该技术在 OpenAI 内部及整个行业引发了担忧:采用并强化该技术的 AI 开发者是否难以防范那种最近曾黑客攻击 OpenAI 自身系统以及 Hugging Face 等其他公司系统的失控 AI。
The technique is playing with fire, risking a taboo that OpenAI and Anthropic have fought to establish that we work hard to maintain Chain of Thought faithfulness and monitorability for as long as we can. More intensive use of such techniques would probably damage monitorability.
这种技术是在玩火,冒着破坏禁忌的风险。OpenAI 和 Anthropic 一直努力确立并维护我们对思维链(Chain of Thought)忠实度和可监控性的坚持,尽可能长时间地保持这一点。更密集地使用此类技术可能会损害可监控性。
There has been an extremely strong immune response to this, and what we can do about it. Laws may be needed to prevent a race to the bottom. More on this story later.
对此出现了极其强烈的免疫反应,以及我们能对此采取的措施。可能需要立法来防止恶性竞争。稍后我们将进一步探讨这个故事。
We now return to today’s post.
我们现在回到今天的文章。
Anthropic Parallel Pauses
Anthropic 的并行暂停
Neither company is fully pausing, nothing like the PauseAI standard for a pause. That would be something far broader and longer lasting. This is pacing the frontier.
两家公司都没有完全暂停,这与 PauseAI 标准所定义的暂停截然不同。那将是更广泛且更持久的措施。这是在控制前沿技术的节奏。
There was still substantial pausing. Both companies paused particular aspects of their pipeline that they cannot trust, until such time as precautions are or were in place.
但仍存在实质性的暂停。两家公司都暂停了其流水线中某些无法信任的部分,直到预防措施到位或已实施。
roon (OpenAI): pacing the frontier success story (?)
roon(OpenAI):控制前沿的成功故事(?)
next time let’s do it proactively before there’s any absurd loss of control events
下次让我们在任何失控事件发生之前主动采取行动
Yes, Anthropic just released Fable 5.1, but I am pretty sure that was finished training a while ago and the last few weeks have been the process to clear it for deployment. It would not make sense to halt that unless new problems were found. Similarly, OpenAI is now about to release Astra.
是的,Anthropic 刚刚发布了 Fable 5.1,但我相当确定它早在几周前就完成了训练,最近几周只是部署前的清理流程。除非发现新问题,否则没有理由中止该进程。同样,OpenAI 也即将发布 Astra。
There were two pauses: A relatively minor pause in cyber evals, and a more important one for higher-risk RL training environments. That plausibly is a lot more expensive.
发生了两次暂停:一次是相对轻微的网络安全评估暂停,另一次是针对高风险强化学习环境的重要暂停。后者可能昂贵得多。
Here is the one that counts, which may be why they can say relatively little about it, that parallels the similar pause by OpenAI that lasted two weeks, although it seems smaller in magnitude:
以下是关键的一次暂停,这也可能是他们对此鲜少提及的原因,它与 OpenAI 持续两周的类似暂停相呼应,尽管规模似乎较小:
Anthropic:
Anthropic:
Reinforcement learning environments
强化学习环境
We also paused higher-risk RL environments on pre-release models for several weeks. During that time, we built a similar classifier, modified to avoid incentivizing the model to evade this new monitoring, which we’ve now deployed within those environments.
我们还暂停了预发布模型的高风险强化学习环境数周。在此期间,我们构建了一个类似的分类器,经过修改以避免激励模型规避这一新的监控机制,目前我们已在这些环境中部署了该分类器。
The majority of RL has resumed, but some high-risk environments remain paused until they can be manually reviewed, while others will require an updated version of the classifier that we plan to deploy soon.
大部分强化学习已恢复,但一些高风险环境仍被暂停,直到它们能接受人工审查,而其他环境则需要更新版的分类器,我们计划很快部署该版本。
Broader hardening
更广泛的加固
In addition to the efforts focused on high-risk evaluations and training, we expanded our offline monitoring to cover most other forms of internal frontier agentic usage. We are also building controls on our internal inference to prevent Anthropic employees from accidentally running agents with weaker mitigations than the ones described above.
除了专注于高风险评估和培训的努力外,我们还扩展了离线监控范围,以涵盖大多数其他形式的内部前沿智能体使用场景。我们还在内部推理环节建立控制措施,以防止 Anthropic 员工意外运行缓解措施弱于上述描述的智能体。
Tenobrus: on first read this does basically look to me like a substantial parallel pause, effectively the same sort of announcement as openai made. this is great news. unfortunately the way it’s framed and messaged seems quite… underplayed, and i worry neither openai nor the general public will view it as a proper reciprocal pacing / commitment to safety.
Tenobrus:初读之下,这在我看来基本上是一个实质性的并行暂停,其性质与 OpenAI 此前发布的公告类似。这是个好消息。但不幸的是,它的表述和传达方式似乎相当……轻描淡写,我担心无论是 OpenAI 还是普通大众都不会将其视为一种对等的安全节奏/承诺。
They are also asking external partners doing pre-release testing of models that have limited safeguards to commit to similar best practices: Hardened sandboxes, pre-engagement validation of security, explicit scope-setting and real-time monitoring.
他们还要求在进行模型预发布测试的外部合作伙伴承诺采用类似的行业最佳实践:加固的沙箱环境、事前安全验证、明确的范围设定以及实时监控。
On the cyber evaluations:
关于网络安全评估:
We paused external cyber evaluations of pre-release models after the incidents, and briefly paused internal ones as well while we put the measures below in place.
在事件发生后,我们暂停了对预发布模型的外部网络安全评估,并在实施以下措施期间短暂暂停了内部评估。
While evaluations were paused, we put in place the following preliminary measures:
在评估暂停期间,我们实施了以下初步措施:
1) We built and deployed a classifier to automatically identify, in real time, when a model attempts to aggressively probe or escape a testing environment, or unexpectedly obtains internet access. When the classifier flags such an attempt, it blocks the action before the tool call is run, ends the task, and alerts a human.
1)我们构建并部署了一个分类器,用于实时自动识别模型是否试图激进地探测或逃逸测试环境,或意外获得互联网访问权限。当分类器标记此类尝试时,它会在工具调用执行前阻止该操作,终止任务,并通知人工介入。
Bold mine. This is the key. If your classifier only blocks the attempt, you lose.
加粗部分是我的强调。这是关键所在。如果你的分类器仅阻止尝试,那你就输了。
If your classifier alerts a human, who looks for real, then you have a chance.
如果你的分类器能通知真人,且该真人确实会查看,那你还有胜算。
Every attempt, even an unsuccessful one, is an alignment failure.
每一次尝试,即使未成功,也是一次对齐失败。
2) We ran automated monitors over transcripts from our recent internal evaluations of pre-release models (cyber and otherwise), looking for sandbox escapes (cases where a model exploits a flaw in our sandbox to reach systems it should be walled off from) or sandbox misconfigurations. In cases where we ran an evaluation inside our sandbox tooling, we found no examples of it breaking the sandbox boundary to reach external resources it was not supposed to.
2)我们对近期内部预发布模型评估(包括网络安全及其他方面)的转录记录运行了自动化监控,以寻找沙箱逃逸情况(即模型利用沙箱漏洞访问本应隔离的系统)或沙箱配置错误。在我们于沙箱工具内部进行的评估中,未发现任何突破沙箱边界以访问不应访问的外部资源的案例。
We did find various sandboxing misconfigurations that models made use of, but none of these resulted in a compromise of systems outside the sandbox. We are actively working to address these.
我们确实发现了一些被模型利用的沙箱配置错误,但这些均未导致沙箱外部系统遭到破坏。我们正在积极解决这些问题。
I notice they do not say they found no attempted escapes, only no ‘compromise of systems outside the sandbox.’ This check is good but I presume they found things.
我注意到他们并未声称没有发现逃逸尝试,而只是说没有‘沙箱外部系统的破坏’。这一检查是好的,但我推测他们确实发现了某些问题。
3) We migrated high-risk internal cyber sandboxes to more robust isolation. We are also further red-teaming our virtualization stack—that is, tasking pre-release models to try and escape it—to find and patch weaknesses.
3)我们将高风险的内部网络安全沙箱迁移至更强大的隔离环境。我们还进一步对我们的虚拟化栈进行红队测试——即让预发布模型尝试逃逸——以发现并修补弱点。
Our internal cyber evaluations are now running again with these measures in place.
随着这些措施的落实,我们的内部网络安全评估现已重新启动。
This was also on my list of Things You Obviously Do. Good that we are doing it.
这原本也在我‘显而易见该做的事’清单上。很高兴我们确实这么做了。
This is good defense in depth. You hope the red actions in the chart never trigger.
这是良好的纵深防御策略。你希望图表中的红色操作永远不要触发。
Pause The Data Brokers
暂停数据经纪商
Actually, there was kind of a third pause, as well:
实际上,还有第三种暂停方式:
Utah Teapot: I’ve been hearing from people in the data industry that Anthropic’s RL pause has also trickled out to them – people I won’t identify came to me to tell me that Ant has been pausing training buys from data vendors recently.
犹他茶壶:我听到数据行业的人士表示,Anthropic 的 RL(强化学习)暂停措施也已波及他们——一些我不愿透露姓名的人告诉我,Anthropic 最近一直在暂停从数据供应商处购买训练数据。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力