跳到主内容
@wquguru
精选75The Zvi(RSS)行业动态

前沿节奏之争:Pacing 信函引发讨论

The Pacing of the Frontier

原文
发到 X

In the wake of the letter calling on us to prepare to potentially Pace the Frontier, there has been much discussion of when pacing the frontier would be prudent, and whether it makes sense to prepare to do so.

在收到那封呼吁我们准备可能“跨越前沿”的信之后,关于何时跨越前沿才是明智的,以及为此做准备是否有意义,人们进行了大量讨论。

This has now been informed by the events surrounding OpenAI training models for months while they had access to a joint de facto message board, which was detected only in the wake of the hacking of HuggingFace by OpenAI’s AIs models during a cybersecurity eval. As we find out more about that, a lot of people have grown far more alarmed, as they should given what they previously believed about the difficulty of alignment, about the state of capabilities and about the level of operational supervision, infrastructure, safety and safety culture at the frontier labs.

现在,这一讨论因围绕OpenAI的事件而变得更加深入:OpenAI在数月内训练模型,期间它们可以访问一个联合的事实上的留言板,这一情况直到OpenAI的AI模型在一次网络安全评估中入侵HuggingFace之后才被发现。随着我们对这件事了解得更多,许多人变得更加警觉,这是理所当然的,因为他们之前对对齐的难度、能力水平以及前沿实验室的运营监督、基础设施、安全性和安全文化所持有的看法,现在都受到了挑战。

This post will not go further into the details of that incident. It treats that as background to keep in mind, and mostly involves perspectives from before the Black Hat talk. This was originally scheduled for Friday and got bumped.

这篇文章不会深入探讨该事件的细节,而是将其作为背景信息,主要涉及Black Hat演讲之前的观点。这篇文章原定于周五发布,但被推迟了。

A lot of the disagreements about the need to pace tie into expectations about the default pace of capability advancements. As I wrote recently in The Three AI Pills, sincere disagreements about AI policy usually boil down to disagreements about the expected pace of progress, and what we expect future AIs will be able to do.

关于是否需要“跨越”的许多分歧,与对能力进步默认速度的预期有关。正如我最近在《三颗AI药丸》中所写,关于AI政策的真诚分歧通常归结为对进步速度的预期差异,以及我们期望未来AI能做什么。

Table of Contents

目录

  • Danger, Will Robinson.
  • Progress Fast and Slow.
  • Statements of Support For Pacing the Frontier.
  • No One In Charge.
  • Pacing The Frontier.
  • Pausing the Frontier.
  • Senator Sanders Demands A Pause.
  • Moderate Prudence.
  • That Escalated Quickly.
  • If You Are In Mundane Alignment Pivot To Scalable Alignment.
  • Taking It Fast.
  • Full Speed Ahead.
  • Suicide Squad.
  • Prepare To Adjust Your Pace.
  • 危险,威尔·罗宾逊。
  • 进步的快与慢。
  • 支持跨越前沿的声明。
  • 无人负责。
  • 跨越前沿。
  • 暂停前沿。
  • 参议员桑德斯要求暂停。
  • 适度谨慎。
  • 事态升级迅速。
  • 如果你从事平凡的对齐工作,转向可扩展对齐。
  • 快速推进。
  • 全速前进。
  • 自杀小队。
  • 准备调整你的步伐。

Danger, Will Robinson

危险,威尔·罗宾逊

There are also disagreements about how dangerous a given level of capability would be, or how it would physically impact the world, and disagreements about what options we have, the nature of various coordination mechanisms or government interventions, and balancing different sacred values. Often people only properly see one half of a key trade-off.

对于特定能力水平有多危险,或它将如何实际影响世界,也存在分歧;对于我们可以选择的方案、各种协调机制或政府干预的性质,以及如何平衡不同的神圣价值,也存在分歧。通常人们只能正确看到关键权衡的一半。

Mostly the reason people often sincerely only see one half of such key tradeoffs is that they anticipate so little AI progress that mitigating existential risks or worrying about humans losing control is unnecessary.

人们常常真诚地只看到这类关键权衡的一半,主要是因为他们对AI进步的预期如此之低,以至于认为减轻存在性风险或担心人类失去控制是不必要的。

Or they anticipate so much AI progress that mitigating existential and catastrophic risks and maintaining human control has to be the priority, and that necessarily is going to mean some group of people collectively choosing and charting, in some way, a deliberate path through causal space towards outcomes that allow us to survive.

或者他们预见到AI进步如此之快,以至于减轻存在性和灾难性风险、保持人类控制必须成为优先事项,这必然意味着某个群体以某种方式共同选择并规划一条穿越因果空间的深思熟虑的路径,以达成让我们得以生存的结果。

Yes, that necessarily means enabling some group of people to have some collective mechanism to chart some aspects of Earth’s path through causal space, and yes there are reasons to worry about that, but that is why we should work to figure out the best way to do that.

是的,这必然意味着让某个群体拥有某种集体机制来规划地球穿越因果空间的某些方面,而且确实有理由对此担忧,但这就是为什么我们应该努力找出最佳方式。

Progress Fast and Slow

快与慢的进步

Those opposing the Pacing the Frontier letter do not want to ‘slow down’ and instead want AI to go ‘fast,’ but their vision of fast is, while super fast by historic standards, not all that fast.

那些反对《为前沿定速》公开信的人不想“放慢速度”,而是希望AI“快进”,但他们眼中的快,虽然按历史标准是超快的,但并没有那么快。

Those who signed the Pacing the Frontier letter mostly want AI to go at least as fast as the opposition. What they want to avoid is AI going super duper ultra hyper fast.

签署《为前沿定速》公开信的人大多希望AI至少和反对者一样快。他们想要避免的是AI变得超级超级超级快。

Daniel Eth (AI Safety): Here’s how I’m interpreting the words:

Daniel Eth(AI安全):以下是我对这些词的理解:

Pause: step on the brakes

暂停:踩刹车

Pace: don’t attach a giant fucking rocket engine on the back of the car that will accelerate us from 65mph to 10,000 mph, at least not unless we can turn it off. also, don’t disable the brakes.

定速:不要在车后安装一个巨大的火箭发动机,将我们从65英里/小时加速到10,000英里/小时,至少除非我们能关掉它。另外,不要禁用刹车。

Almost everyone signing the letter wants better chatbots and doctors and science and other forms of diffusion. They don’t want superintelligence and a singularity in 2027.

几乎每个签署这封信的人都想要更好的聊天机器人、医生、科学和其他形式的扩散。他们不想要2027年的超级智能和奇点。

Nick: a lot of the anti crowd wants better chatbots and doctors and stuff, a lot of the pro crowd expects like way crazier worlds in the short term and wants them to come in the just slightly less short term when we’ve figured out how to control these things better. plenty of overlap

Nick:很多反对者想要更好的聊天机器人和医生之类的东西,很多支持者则期望短期内出现更疯狂的世界,并希望这些世界在稍长一点的时间内到来,那时我们已经更好地掌握了如何控制这些东西。有很多重叠之处。

how to do it no idea, and also how to measure the speed limit no idea, but i feel like this framing avoids some of the issues with pause, which also has roughly the same questions

怎么做我不知道,如何测量速度限制我也不知道,但我觉得这种框架避免了一些暂停的问题,暂停也有大致相同的问题。

Dean W. Ball: This is what most people I know who signed the “pacing” letter (myself included) think. The slowdown we have in mind is temporary, and to a rate of progress that is still much faster than even today’s rate.

Dean W. Ball:这是我所认识的大多数签署“定速”信的人(包括我自己)的想法。我们设想的放缓是暂时的,而且放缓后的进步速度仍然比今天快得多。

Augmented Fifth: It’s Dec 19, school is out, and the Christmas presents are under the tree. “Let’s wait until February” is not going to fly with the kids.

Augmented Fifth:现在是12月19日,学校放假了,圣诞礼物在树下。“等到二月再说”对孩子来说可不行。

I quote that last one because what is happening is that the people who are arguing against the Pacing the Frontier letter are saying they won’t wait until February and demanding they get the presents on December 25 and they’d better get that BB gun.

我引用最后那句话,是因为现在反对《Pacing the Frontier》公开信的人说,他们不会等到二月,要求必须在12月25日收到礼物,而且最好能拿到那把BB枪。

Whereas those signing the letter are saying maybe we should wait until at least tomorrow before we order even more presents and there is no more room left in the house, plus maybe keep the toys reasonable, and only get you the BB gun that’ll shoot your eye out kid and not an AK-47 or tactical nuke or that sexy Von Neumann probe.

而那些签署公开信的人则说,也许我们应该至少等到明天再订购更多礼物,因为家里已经没有更多空间了,而且也许应该让玩具保持合理,只给你那把会打掉你眼睛的BB枪,而不是AK-47、战术核弹或那个性感的冯·诺依曼探测器。

Statements of Support For Pacing the Frontier

对《Pacing the Frontier》的支持声明

Another strong comment from one of those who signed the Pacing the Frontier letter:

另一位签署《Pacing the Frontier》公开信的人发表了强有力的评论:

Drake Thomas (Anthropic): Very happy to have signed the pacing the frontier letter; its existence gives me a lot of hope for humanity’s survival!

Drake Thomas(Anthropic):非常高兴签署了《Pacing the Frontier》公开信;它的存在让我对人类生存充满希望!

I endorse the letter as written (and think, given the constraints, it’s probably close to the best it could be for this level of consensus), but some places where I differ from its connotational tone:

我赞同这封信的书面内容(并且认为,考虑到各种限制,对于这种共识水平来说,它可能已经接近最佳状态了),但在某些方面,我与它的隐含语气有所不同:

(1) Not only is AI “not guaranteed” to make a dramatically better future, the odds of failure are terrifyingly high: I think* there’s something like a 40% chance we get an outcome around as bad as human extinction or worse, and another 30% chance we get a future that, while containing some good things, falls radically short of what a wiser civilization could have obtained (say, <5% of the value of a truly great future).

(1)人工智能不仅“不能保证”带来更美好的未来,失败的可能性高得可怕:我认为*我们大约有40%的概率会得到一个与人类灭绝一样糟糕或更糟的结果,还有30%的概率会得到一个虽然包含一些美好事物,但远未达到一个更明智文明本可以获得的未来(比如,不到一个真正伟大未来价值的5%)。

(2) I don’t really endorse the vibes of “to realize AI’s potential”; I think a much more immediate and pressing motivation is “to avoid catastrophically bad outcomes from AI that will kill a lot of people”. (The potential is also very important, ofc, but I think most moral theories would view it as being of secondary importance when risks are this high and there’s little that would do more than temporarily delay that potential anyway.)

(2)我并不完全赞同“实现人工智能潜力”这种氛围;我认为一个更直接、更紧迫的动机是“避免人工智能带来的灾难性后果,这些后果将导致大量人员死亡”。(潜力当然也非常重要,但我认为大多数道德理论会认为,当风险如此之高,而且几乎没有什么能比暂时推迟这种潜力更有用时,潜力是次要的。)

(3) “address emerging risks, develop security measures, and strengthen oversight” is fine so far as it goes but a little vague. Concrete things I’d like to do with slower AI development: way better interpretability, build up a robust and well-resourced third party ecosystem for independent auditing of AI companies and get lots of reps in for their oversight, put tons of effort into the automation of alignment research, work on governance mechanisms for ASI, develop a vastly better science of the nature and development of AI behavior, build much more powerful control mechanisms, point lots of powerful AI labor at ambitious scalable alignment projects (eg work like ARC’s), deep dives (including external audits) of individual AI behavior incidents, etc. Also getting civilizational biosecurity preparedness in order.

(3) “应对新兴风险、制定安全措施并加强监管” 就目前而言尚可,但略显模糊。在减缓人工智能发展方面,我希望具体做的事情包括:大幅提高可解释性,建立强大且资源充足的第三方生态系统,用于对人工智能公司进行独立审计,并为其监管积累大量经验;投入大量精力实现对齐研究的自动化;致力于 ASI 的治理机制;发展一门关于人工智能行为本质和发展的更完善科学;构建更强大的控制机制;将大量强大的人工智能劳动力投入到雄心勃勃的可扩展对齐项目中(例如 ARC 的工作);对个别人工智能行为事件进行深入调查(包括外部审计)等。此外,还要做好文明层面的生物安全防范准备。

*epistemic status very approximate vibes, my numbers will change day to day and depending on the exact operationalization.

*认知状态:非常粗略的感觉,我的数字会随日期和具体操作化方式而变化。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近