Anthropic 八月风险报告解读:内部模型 Model 2 超
Anthropic Risk Report: August 2026
I am grateful that Anthropic is producing periodic Risk Reports.
我很感激Anthropic定期发布风险报告。
At first I was skeptical. It turns out I was wrong. Anthropic is revealing a lot of new information, some of it rather alarming, that it did not have to disclose, and is providing detailed insight into how they think about things. This is very cool.
起初我持怀疑态度。结果证明我错了。Anthropic披露了大量新信息,其中一些相当令人担忧,这些信息本不必公开,而且他们提供了对其思考方式的详细见解。这非常酷。
Thus I found this report to be a moderately positive update overall, if we presume they are not silently omitting the worst of it. There are a bunch of not great things we find out about, but I would have expected some set of mistakes at least as bad, and I wouldn’t have expected them to choose to tell us about all of it.
因此,如果假设他们没有默默隐瞒最糟糕的部分,我认为这份报告总体上是一个适度积极的更新。我们发现了一些不太好的事情,但我原本预期会有一系列至少同样糟糕的错误,而且我没想到他们会选择告诉我们所有这一切。
It does mean one more set of 186 page documents I have to read every so often, almost all of which is meaningfully new material this time around.
这确实意味着我每隔一段时间就得再读一份186页的文件,而这次几乎全部内容都是有意义的新材料。
The other revelation is the existence of the world’s likely best model, ‘Model 2.’
另一个揭示是可能存在世界上最好的模型——‘模型2’的存在。
This was a rough one to fully get through, so apologies in advance for any errors of interpretation.
这份报告读起来相当艰难,所以提前为任何解读错误道歉。
Table of Contents
目录
- Agent Model 1 and Agent Model 2.
- Executive Summary (1).
- The Rules Are Serious But Not Literal.
- Misalignment Is a State of Mind (2.5).
- Autonomy Threat Model 1: Misalignment in High-Stakes Settings (2).
- Some Strange Uses Of The Word Safe I Wasn’t Previously Aware Of.
- Now Versus Future (2.17).
- The Core Claims And Argument (2.6).
- The Rest of the Important Arguments In Section 2.
- Risk Assessment (2.19).
- Pre-Internal-Deployment Review (2.18).
- A Guide To Internal Use Monitoring (2.23.1).
- Blocking Interventions (2.23.2).
- The Power Seeking Environment Evaluation (2.24).
- Opus 4.8-Reward-Hacker (2.25).
- Autonomy threat model 2: Risks from automated R&D (3).
- Yes That Does Seem Kind Of Risky.
- Could We Replace Our Researchers?
- How Much Could We Be Accelerating Our AI Researchers?
- What Could Possibly Go Wrong If We Replaced Our Researchers?
- Risk Mitigations For AI R&D Automation.
- Overall Risk From Automation of AI R&D.
- Biological and Technically Also Chemical Weapons Production.
- The Threat Models for Biological and Chemical Weapons.
- Model Capabilities (4.4).
- Classifiers (4.5).
- Acceleration Dynamics (5.1).
- Distillation (5.1.1).
- Safety Process Failures (5.2).
- Refusing To Find Innovative Misalignment Techniques (5.2.2).
- Exposing the Chain of Thought Reasoning To Grading Pressure Quite a Lot (5.2.3).
- Directly Training On Misaligned Behavior During a Production Training Run (5.2.4).
- An instance of unmonitored unrestricted agents with access to sensitive resources (5.2.5).
- Repeated training on alignment-faking transcript datasets (5.2.6).
- Benefits From Anthropic’s Operating as a Frontier AI company (5.3).
- Model Weight Security (6.4).
- Risk Has Been Reported.
- 代理模型1和代理模型2。
- 执行摘要(1)。
- 规则是严肃的,但不是字面上的。
- 错位是一种心态(2.5)。
- 自主性威胁模型1:高风险环境中的错位(2)。
- 一些我以前不知道的‘安全’一词的奇怪用法。
- 现在与未来(2.17)。
- 核心主张和论证(2.6)。
- 第2节中其余的重要论证。
- 风险评估(2.19)。
- 内部部署前审查(2.18)。
- 内部使用监控指南(2.23.1)。
- 阻断干预(2.23.2)。
- 权力寻求环境评估(2.24)。
- Opus 4.8-奖励黑客(2.25)。
- 自主性威胁模型2:自动化研发的风险(3)。
- 是的,那看起来确实有点冒险。
- 我们能否替换我们的研究人员?
- 我们能让AI研究人员加速多少?
- 如果我们替换研究人员,可能会出什么问题?
- AI研发自动化的风险缓解措施。
- AI研发自动化的整体风险。
- 生物以及技术上还有化学武器的生产。
- 生物和化学武器的威胁模型。
- 模型能力(4.4)。
- 分类器(4.5)。
- 加速动态(5.1)。
- 蒸馏(5.1.1)。
- 安全流程失败(5.2)。
- 拒绝寻找创新性的错位技术(5.2.2)。
- 在评分压力下大量暴露思维链推理(5.2.3)。
- 在生产训练运行期间直接对错位行为进行训练(5.2.4)。
- 一个未受监控、无限制且能访问敏感资源的代理实例(5.2.5)。
- 在一致性伪造转录数据集上重复训练(5.2.6)。
- Anthropic作为前沿AI公司运营的益处(5.3)。
- 模型权重安全(6.4)。
- 风险已被报告。
Agent Model 1 and Agent Model 2
代理模型1和代理模型2
There are two internal models referenced in this report.
本报告中提到了两个内部模型。
Model 1 is broadly similar to Mythos Preview and Mythos 5, and is not expected to have wide deployment, externally or internally.
模型1大致类似于Mythos Preview和Mythos 5,预计不会在外部或内部广泛部署。
Model 2 is ‘somewhat more capable than Mythos 5,’ a noticeable improvement for internal use, but ‘does not display a capability jump of the degree observed from Claude Opus 4.6 to Mythos Preview.’ This model is going to be internal-use only.
模型2“比Mythos 5能力更强一些”,对于内部使用来说是一个显著的改进,但“没有展现出从Claude Opus 4.6到Mythos Preview所观察到的能力跳跃程度”。该模型将仅供内部使用。
The jump from Opus 4.6 to Mythos Preview was big. Several release cycles big. We get a lot more data throughout the report, but given Model 2 gets that comparison point, one would presume that for internal purposes Model 2 is multiple releases ahead.
从 Opus 4.6 到 Mythos Preview 的跳跃是巨大的,跨越了好几个发布周期。我们在整个报告中获得了更多数据,但鉴于 Model 2 获得了那个对比点,可以推测在内部用途上,Model 2 领先了多个版本。
To contrast with that, on the AECI it is only 1.5 points ahead, which would be only about a month of progress. My guess is this is a low estimate, and that Model 2 was intentionally not decked out with a bunch of general skills given its purpose. Whereas on the test of substituting for Anthropic’s researchers we see a big jump from 54.8% (Mythos Preview) and 50.3% (Mythos 5) to 62.8% for Model 2.
相比之下,在 AECI 上它仅领先 1.5 分,这大约只相当于一个月的进展。我猜测这是一个低估,并且 Model 2 因其用途而未特意配备大量通用技能。而在替代 Anthropic 研究人员的测试中,我们看到从 54.8%(Mythos Preview)和 50.3%(Mythos 5)大幅跃升至 Model 2 的 62.8%。
There are multiple plausible good reasons for not releasing Model 2. Model 2 could be highly specialized for internal tasks. It could also be too capable or dangerous to release, either due to harm or the worry it would accelerate others. The relatively low adoption rate for Fable 5 points towards wanting to hold such a model back.
不发布 Model 2 可能有多个合理的理由。Model 2 可能高度专用于内部任务。它也可能过于强大或危险而不宜发布,无论是出于危害考虑还是担心它会加速他人进展。Fable 5 相对较低的采用率表明有意保留这样的模型。
Executive Summary (1)
执行摘要(1)
The report is divided by each type of risk, by threat model, by relevant AI models, and details current capabilities, behaviors, mitigations and overall level of risk. It then looks forward to the future and offers industry-wide recommendations.
报告按每种风险类型、威胁模型、相关 AI 模型划分,并详细说明了当前能力、行为、缓解措施和总体风险水平。然后展望未来并提供行业范围的建议。
Risk reports cover events up to a month prior, so the coverage date is July 15, 2026. A year ago that would have seemed fine. Now it seems like quite a while, huh?
风险报告涵盖截至一个月前的事件,因此覆盖日期为 2026 年 7 月 15 日。一年前这似乎没问题。现在感觉已经过了一段时间,对吧?
The Long-Term Benefit Trust (TLBT) is authorized to request external review of the report, but has not done so. I would at minimum do so next time if I was them.
长期利益信托(TLBT)被授权要求对报告进行外部审查,但尚未这样做。如果我是他们,我至少下次会这样做。
I like that they cover both ‘what is the absolute risk from this’ and ‘what is the marginal risk from this given everyone else exists.’
我喜欢他们同时涵盖‘这带来的绝对风险是什么’和‘考虑到其他所有人的存在,这带来的边际风险是什么’。
They deal with autonomy, automated AI R&D and biological and chemical weapons production as risks. It is odd, even now, to exclude cyber risks from the core threat models here. I understand the argument from insufficiently catastrophic on its own, but in practice I think you need to deal with it.
他们将自主性、自动化 AI 研发以及生物和化学武器生产视为风险。即使到现在,将网络风险排除在核心威胁模型之外也显得奇怪。我理解其自身不够灾难性的论点,但在实践中我认为你需要处理它。
At several points, the risk report essentially concedes versions of my objections, but then forgets that it conceded them and doesn’t alter its conclusions.
在多个地方,风险报告实质上承认了我反对意见的某些版本,但随后忘记了这些承认,并未改变其结论。
Then in Section 5 they get to the more interesting questions.
然后在第 5 节,他们转向了更有趣的问题。
The Rules Are Serious But Not Literal
规则是严肃的,但不是字面上的
The risk threshold definitions have changed from simpler statements to ones with more specific detail.
风险阈值的定义已从更简单的陈述变为更具体详细的陈述。
For AI R&D automation I think the new wording is harder to parse but fine. For novel biological weapons production, the new version narrows to only look at ‘substitute for the scarce human expertise’ to the exclusion of other methods. This is a reasonable primary threat model or scenario, but I worry about looking only at that path.
对于AI研发自动化,我认为新措辞更难解析,但尚可接受。对于新型生物武器生产,新版本缩小范围,仅关注‘替代稀缺的人类专业知识’,排除了其他方法。这是一个合理的首要威胁模型或情景,但我担心只关注这一路径。
The both good and bad news is I take such documents seriously, but I no longer take such documents literally, in either direction.
好坏参半的消息是,我认真对待这类文件,但我不再字面理解这类文件,无论方向如何。
- If Anthropic or another frontier lab has a model that technically crosses their risk threshold, but in a way that seems to be harmless, I expect them to modify their threshold or find some other way to ignore this event.
- If Anthropic or another frontier lab has a model that technically does not cross their risk threshold, but is obviously risky in the way the rules tried to measure, I expect them to act as if it had crossed the threshold.
- 如果Anthropic或其他前沿实验室有一个模型在技术上跨越了其风险阈值,但方式看似无害,我预计他们会修改阈值或寻找其他方式忽略此事件。
- 如果Anthropic或其他前沿实验室有一个模型在技术上未跨越其风险阈值,但显然在规则试图衡量的方式上存在风险,我预计他们会像跨越阈值一样行动。
As always: I would like to see thresholds honored even in cases where you think it ‘should not count,’ because that is what commitments are for, and you should expect rationalizations to occur as you rush to make releases, and you need to set a good example on this. The whole original idea was if-then commitments, where if [X] happens you do [Y], agreed upon in advance, but our civilization seems to lack this technology. We can’t, we don’t know how.
一如既往:我希望即使在你们认为‘不应计入’的情况下也能遵守阈值,因为这就是承诺的意义,你们应该预期在急于发布时会出现合理化解释,并且需要在这方面树立良好榜样。最初的想法是如果-那么承诺,即如果发生[X]则执行[Y],事先达成一致,但我们的文明似乎缺乏这种技术。我们不能,我们不知道如何做到。
I’m not saying you would never modify your rules to be more lenient, if you have good evidence that you should do that. If you can never make the rules more lenient then the rules have to never be strict. We don’t want that. But the bar for changing the rules, especially in response to a pending event, needs to be a lot higher than it is, and we need to be willing to endure some actual costs even when we now think it is dumb.
我并不是说如果你们有充分证据表明应该放宽规则,你们永远不会修改规则。如果规则永远不能放宽,那么规则就必须永远不严格。我们不希望那样。但修改规则的障碍,尤其是针对即将发生的事件,需要比现在高得多,而且我们需要愿意承受一些实际成本,即使我们现在认为这很愚蠢。
For now, Anthropic and also OpenAI have made what seem like good decisions in terms of responding to threats and deciding what to release and not release when, based on what is known at the time, even if they make mistakes elsewhere. That is good. I hope it continues.
目前,Anthropic和OpenAI在应对威胁以及根据当时已知情况决定发布什么、何时发布方面做出了看似正确的决策,即使他们在其他方面犯了错误。这是好的。我希望这种情况能持续下去。
Misalignment Is a State of Mind (2.5)
错位是一种心态(2.5)
The definition section is greatly appreciated.
定义部分非常值得赞赏。
I like this definition a lot, although I have important nitpicks:
我非常喜欢这个定义,尽管我有重要的吹毛求疵之处:
Misalignment [definition]: Misalignment is a latent property of a specific computation performed by a model in a given context.
错位[定义]:错位是模型在特定上下文中执行的特定计算的潜在属性。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力