跳到主内容
@wquguru
精选88The Zvi(RSS)行业动态多源精选 ×7

Anthropic提出前沿AI节奏控制方案获OpenAI等支持

We Must Pace The Frontier

原文
发到 X
推荐理由

AI行业头部公司首次就“主动放缓研发节奏”达成广泛共识,嵌入式评估机制若落地将重塑行业安全标准,值得从业者密切关注后续执行细节。

Dario Amodei has a new essay that finally says the thing: We Must Pace the Frontier, naming his call after the Pacing the Frontier letter lab employees signed in July.

达里奥·阿莫迪(Dario Amodei)发表了一篇新文章,终于说出了那句话:我们必须为前沿技术设定节奏。他将自己的呼吁命名为“为前沿技术设定节奏”,呼应了七月时各大实验室员工签署的《为前沿技术设定节奏》公开信。

As in, we need to slow the rate at which AIs increase their capabilities, to allow for the necessary alignment and safety work.

也就是说,我们需要减缓人工智能提升其能力速度的步伐,以便为必要的对齐和安全工作留出时间。

He explained that, without pacing, he expects things to escalate quickly. He offered three proposals, and unilaterally committed to the first one. OpenAI followed, and both Elon Musk and Demis Hassabis endorsed the overall proposal.

他解释说,如果不进行节奏控制,他预计事态会迅速升级。他提出了三项建议,并单方面承诺执行第一项。OpenAI 随后跟进,埃隆·马斯克(Elon Musk)和杰米斯·哈萨比斯(Demis Hassabis)都支持这一整体提议。

There is still a long way to go. The odds are still against us. The situation remains grim. The hard part lies ahead. We do not agree on what ‘Pacing the Frontier’ will mean in practice. But this is Actual Progress. The work can begin.

前路依然漫长。胜算仍在我们这边不利。局势依然严峻。艰难的部分还在后面。我们对于“为前沿技术设定节奏”在实践中意味着什么尚未达成一致。但这是真正的进展。工作可以开始了。

Table of Contents

目录

  • Pacing Does Not Mean Pausing.
  • Dario’s First Proposal: Embedded Evaluators.
  • Dario’s Second Proposal: Democratic Coordination.
  • Dario’s Third Proposal: Global Coordination.
  • Why Pace Now?
  • Sam Altman Agrees and Commits to Embedded Evaluators.
  • OpenAI Will Not IPO This Year.
  • Elon Musk Agrees.
  • Demis Hassabis Agrees.
  • Microsoft CEO Satya Nadella Agrees And Talks His Book.
  • Anthropic’s Long-Term Benefit Trust Is On Board.
  • General Online Reactions.
  • Mainstream Press Coverage.
  • OpenAI Researcher Explains What The Labs See And It’s a Rocket Ship.
  • Consider the Alternative.
  • It’s Totalitarianism, Joe.
  • Yes We’re The Baddies How Did You Know?
  • Sometimes People On the Internet Just Lie.
  • David Sacks Groks The Situation.
  • David Sacks Says Go Ahead.
  • Lies and Confusions About Who Previously Claimed What.
  • House Speaker Mike Johnson Wants To Lock Everyone In a Room.
  • Donald Trump is Not Tired of Winning.
  • We Must Avoid Polarization on AI at (Almost) All Costs.
  • How Will We Know If They Actually Paced?
  • The Real Frontier Is Internal Models At Top Labs.
  • 设定节奏并不意味着暂停。
  • 达里奥的第一项提议:嵌入式评估员。
  • 达里奥的第二项提议:民主协调。
  • 达里奥的第三项提议:全球协调。
  • 为什么现在就要设定节奏?
  • 山姆·阿尔特曼同意并承诺实施嵌入式评估员。
  • OpenAI 今年不会上市。
  • 埃隆·马斯克表示同意。
  • 杰米斯·哈萨比斯表示同意。
  • 微软首席执行官萨提亚·纳德拉表示同意,并谈论他的新书。
  • Anthropic 的长期利益信托基金加入其中。
  • 网络上的普遍反应。
  • 主流媒体报道。
  • OpenAI 研究人员解释了实验室看到的情况,那简直是一艘火箭飞船。
  • 考虑一下替代方案。
  • 那是极权主义,乔。
  • 是的,我们是坏人,你是怎么知道的?
  • 有时网上的人只是在撒谎。
  • 大卫·萨克斯洞悉了这一局势。
  • 大卫·萨克斯说:放手去做。
  • 关于谁此前声称什么的谎言与混淆。
  • 众议院议长迈克·约翰逊想把所有人锁在一个房间里。
  • 唐纳德·特朗普并不厌倦胜利。
  • 我们必须不惜(几乎)一切代价避免在人工智能问题上两极分化。
  • 我们如何知道他们是否真的放慢了步伐?
  • 真正的边疆在于顶级实验室的内部模型。

Pacing Does Not Mean Pausing

控制节奏并不意味着暂停

AI capabilities are improving very fast. Even I cannot keep up.

人工智能能力正在飞速提升。就连我也跟不上。

Think about what models were like even one year ago.

想想一年前的模型是什么样的。

Inside Anthropic and OpenAI, internal models are improving even faster. Over the summer there was a step change, as Mythos and Astra started kicking off the early stages of recursive self-improvement (RSI).

在 Anthropic 和 OpenAI 内部,内部模型的改进速度甚至更快。今年夏天出现了一个质的飞跃,因为 Mythos 和 Astra 开始启动递归自我改进(RSI)的早期阶段。

Dario Amodei: My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all.

达里奥·阿莫代伊:我的第一个担忧是,大约从今年夏天开始,人工智能的发展速度急剧加快,这主要是由人工智能日益增强的构建下一代人工智能的能力所驱动的。这种动态被称为递归自我改进,它已经开始在整个行业发生,包括在 Anthropic,正如我们和其他人所描述的那样。如果不加限制,它可能会超出我们理解和控制这些系统的能力,因此必须非常谨慎地推进,甚至可能根本不应该推进。

Dario worries that we by default are 6-12 months away from a rogue swarm of AI agents, similarly misaligned to the ones in the OpenAI-HuggingFace incident, being able to use a persistent botnet to take over the internet, or worse.

达里奥担心,按照默认情况,我们距离失控的人工智能代理群只有 6-12 个月的时间,这些代理群类似于 OpenAI-HuggingFace 事件中出现的那些未对齐的代理群,能够利用持久的僵尸网络接管互联网,或者更糟的情况。

That is how fast he expects default progress to be. If we went a lot less fast than that, it would still be extremely fast.

这就是他对默认进展速度的预期。如果我们比这慢得多,那仍然会非常快。

Thus (bold his):

因此(他加了粗体):

Dario Amodei: We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain.

达里奥·阿莫代伊:我们必须减缓提高人工智能模型能力的步伐。进步看起来仍然会很快,我们必须明智地利用我们争取到的时间。

Pacing does not mean pausing. As Dario Amodei says, progress will still seem fast.

控制节奏并不意味着暂停。正如达里奥·阿莫代伊所说,进步看起来仍然会很快。

I strongly encourage you to read the original essay in its entirety.

我强烈建议你完整阅读这篇原始文章。

Dario’s First Proposal: Embedded Evaluators

达里奥的第一个提议:嵌入式评估器

This is an excellent proposal. I am very happy that Anthropic and OpenAI will do it.

这是一个极好的提议。我很高兴 Anthropic 和 OpenAI 会这样做。

If we do not know what is going on inside the labs, we cannot do anything about it, and the labs have to worry that their competitors are speeding ahead.

如果我们不知道实验室内部发生了什么,我们就无法采取任何措施,而且实验室不得不担心他们的竞争对手正在加速前进。

Dario frames the benefits as:

达里奥将好处表述为:

  • Verifiability: You can check to see if the rules are being followed.
  • Transparency: The public can have a better idea what the hell is going on.
  • Second Opinion: Having an informed opinion free of commercial incentives.
  • 可验证性:你可以检查规则是否得到遵守。
  • 透明度:公众可以更好地弄清楚到底发生了什么。
  • 第二项意见:提供不受商业利益影响的知情观点。

Thus the first proposal enables the second and third proposals. It is also a good idea anyway. We need more visibility into the labs. It would have been very good to have such evaluators during recent incidents. So I call upon the other major labs, that have not yet done so, to also commit to this first step.

因此,第一项提案使第二项和第三项提案成为可能。无论如何,这本身也是一个好主意。我们需要对实验室有更高的透明度。在最近的事件中,如果有这样的评估者将会非常有益。因此,我呼吁其他尚未采取此行动的主要实验室也承诺迈出这一步。

I agree with Dario that this should be made mandatory in its full form. If you are well-resourced enough to pursue plausibly frontier models, then you can afford to do this, especially if you are our biggest open model advocates, as in Meta, Google or Nvidia.

我同意达里奥的观点,即应以完整形式强制实施这一要求。如果你有足够的资源去追求合理的前沿模型,那么你就有能力做到这一点,尤其是像Meta、Google或Nvidia这样我们最大的开源模型倡导者。

If you think that your operation could not survive if there were embedded evaluators checking its safety, then ask yourself why you think this.

如果你认为你的运营在存在嵌入式评估人员检查其安全性时无法生存,那么请自问为何会有这种想法。

Embedded Evaluators. Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes.

嵌入式评估者。每家前沿AI公司都承诺向一组嵌入式第三方评估者(如METR)提供持续的、类似员工的访问权限,他们的职责是验证对安全实践和承诺的遵守情况,报告事件,并帮助评估不仅包括已完成的AI模型,还包括训练管道和流程的对齐情况。

This is the key step for verifiability of any pacing commitments, and has precedent in the banking industry, which sometimes involves regulatory “supervisors” embedded along with employees.

这是任何节奏承诺可验证性的关键步骤,银行业已有先例,有时监管机构会嵌入“监督员”与员工一起工作。

Anthropic is unilaterally committing to this step now. We intend this to be part of a broader push to redouble efforts on our safety and alignment work.

Anthropic现在单方面承诺采取这一步骤。我们打算将其作为更广泛努力的一部分,以加倍投入安全和对齐工作。

Anthropic is proposing to empower the evaluators quite a bit:

Anthropic提议赋予评估者相当大的权力:

  • Desks in our offices, access badges, and company laptops.
  • Access to workspaces, tools, and permissions mostly comparable to what internal risk assessment teams have. We’ll make some exceptions, such as where the law or our contracts require it, or to protect customers’ and partners’ private information. We’ll also establish strong internal norms reinforcing reviewers’ access to relevant information, including through live conversations with employees.
  • A contract that balances the complexities mentioned above. External reviewers should have the right to publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive — without editorial control by Anthropic. We will have the narrow ability to redact security-sensitive, legally privileged, commercially sensitive, or third-party confidential information, but we can’t redact findings just because they are unfavorable. The reviewers can say publicly if a redaction removed something important to their conclusions.​
  • 在我们的办公室设有工位、门禁卡以及公司笔记本电脑。
  • 主要与内部风险评估团队相当的办公空间、工具和权限访问。我们将做一些例外处理,例如法律或合同要求的情况,或为了保护客户和合作伙伴的私人信息。我们还将建立强有力的内部规范,强化审查员获取相关信息的机会,包括通过与员工进行实时对话。
  • 一份平衡上述复杂性的合同。外部审查员应有权发布有关风险水平、事件、实践以及他们获得或未获得的访问权限的关键发现——不受 Anthropic 的编辑控制。我们将拥有有限的权利来删除安全敏感、法律特权、商业敏感或第三方保密信息,但不能仅仅因为发现结果不利就对其进行删减。如果删减影响了他们结论中的重要内容,审查员可以公开说明这一点。

This is an unusual step for a company, but we think it is important to prove out the concept of embedded external reviewers. Once again, we urge other frontier companies to follow suit.

这对一家公司来说是不寻常的步骤,但我们认为这对于验证嵌入式外部审查员的概念至关重要。我们再次敦促其他前沿公司效仿。

It is easy to imagine a mostly fake version of embedded evaluators. This is promising to very much not be that, and to unusually empower the evaluators to report, if it was fake, that the arrangement was indeed fake. These details, if followed, answer the ‘oh you can just fake this’ objection.

很容易想象出一个几乎虚假的嵌入评估者版本。这有望完全不是那样,并赋予评估者异常强大的权力去报告:如果是假的,那么这种安排确实是假的。如果遵循这些细节,就能回应‘你完全可以伪造这一切’的质疑。

The good objections to this are about implementation. We need enough evaluators, they need to be qualified, and they need to be independent and trustworthy.

对此合理的质疑主要集中在实施层面。我们需要足够多的评估者,他们需要具备资质,并且必须保持独立和值得信赖。

METR is great, but METR cannot do this alone, and we have one hell of a set of incentive problems to solve. We also need to both have them be competent, and also not too linked to the existing ecosystems and labs, and the funding will have to come from somewhere.

METR 很棒,但 METR 无法独自完成这项工作,而且我们有一堆棘手的激励问题需要解决。我们还需要确保他们具备能力,同时不过度依赖现有的生态系统和实验室,资金也必须从某处筹集。

Tim Hwang: The third party AI evaluator can either be independent, knowledgeable, or sustainably funded. Pick two.

Tim Hwang:第三方 AI 评估者要么独立,要么专业,要么资金可持续。三选二。

roon (OpenAI): I pick the latter two. it’s pretty much an impossible ask to find talented ai people who were not in some way in the orbit of the labs in the last decade

roon (OpenAI):我选择后两者。在过去十年里,找到那些在某种程度上未处于实验室影响范围内的优秀 AI 人才几乎是不可能的要求。

David Manheim: I agree that if independence means “never had any contact,” it’s idiotic. But otherwise, it sounds a lot like “no one will bother solving this problem” – sustainable funding via various mechanisms is entirely possible, and we see it occur in other domains. This can be solved.

David Manheim:我同意如果独立性意味着‘从未有过任何接触’,那是愚蠢的。但除此之外,听起来很像‘没人会 bother 解决这个问题’——通过各种机制实现可持续资金是完全可能的,我们在其他领域也看到了这种情况的发生。这个问题是可以解决的。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →