Anthropic研究员辞职警告AI致人类灭绝风险引发行业震荡
Jacob Coxon Warns of Human Extinction and Triggers a Preference Cascade
AI安全领域的标志性事件,多位核心研究员公开质疑超级智能开发节奏,建议关注其引发的行业合规与战略转向。
CEOs of major AI labs, and employees of major AI labs, including OpenAI and Anthropic, often say they plan to build superintelligence soon, as in within a few years create AIs that are superior to humans at essentially all cognitive tasks.
各大AI实验室的首席执行官,以及包括OpenAI和Anthropic在内的各大AI实验室的员工,经常表示他们计划很快构建超级智能,即在几年内创造出在几乎所有认知任务上都优于人类的AI。
They often warn that such AIs might kill everyone. Or that AIs might cause mass unemployment, cause cyberattacks across the internet, enable mass surveillance or risk causing any number of other highly bad things.
他们经常警告说,这样的AI可能会杀死所有人。或者AI可能会导致大规模失业、引发互联网上的网络攻击、实现大规模监控,或导致其他许多极其糟糕的事情的风险。
These warnings are consistently and directly against the interests of the labs. Yet the warnings have recently gotten a lot louder and more frequent. OpenAI has been practically screaming, for those with ears to listen, on many occasions.
这些警告始终且直接违背了实验室的利益。然而,最近这些警告变得响亮得多,频率也高得多。对于那些有耳朵去听的人来说,OpenAI在许多场合几乎是在尖叫。
A series of events, over two months and especially the last week or so, including internal observations of the pace of progress at OpenAI and also Anthropic, have freaked out everyone involved quite a lot more than they were already freaked out.
过去两个月,尤其是最近一周左右的一系列事件,包括对OpenAI和Anthropic内部进展速度的观察,让所有相关人员比以往更加惊恐。
After all the events, plus statements by Dean Ball and Jakub Pachocki, we were already seeing the beginnings of a preference cascade.
在所有这些事件之后,加上Dean Ball和Jakub Pachocki的声明,我们已经看到了偏好级联的开端。
Then along came Jacob Coxon as the tipping point, and things took off.
然后Jacob Coxon作为临界点出现,事情开始爆发。
Table of Contents
目录
- Jacob Coxon Resigns From Anthropic In Protest And Sounds The Alarm.
- Mainstream Media Finally Pays Attention.
- Preference Cascade at Anthropic.
- Preference Cascade at OpenAI.
- Preference Cascade at Google.
- #NotAllMembersOfTechnicalStaff.
- Why a Preference Cascade Now?
- This Is What Many Anthropic and OpenAI Employees Actually Believe.
- To Quit Or Not To Quit.
- Quiet Quitting Is A Dominated Option.
- When You Quit, Very Serious People Understand What That Means.
- Jacob Coxon Believes Existential Risk Is High That Is Why He Quit.
- Evan Hubinger Believes Existential Risk Is High That Is Why He Stays.
- Anthropic and OpenAI Have Commercial Incentives To Downplay Existential Risks, Not Advertise Them.
- What Do We Do Now?
- OK, But How Exactly Would AI Kill Everyone?
- Best Start Believing In Science Fiction Stories Because You Are In One.
- Literal Extinction Is Not Much Harder Than Loss of Control.
- Conspiracytown Is Always Hiring.
- Now You See It.
- Jacob Coxon辞职抗议并敲响警钟。
- 主流媒体终于开始关注。
- Anthropic的偏好级联。
- OpenAI的偏好级联。
- Google的偏好级联。
- #并非所有技术人员都这样认为。
- 为什么现在会出现偏好级联?
- 这是许多Anthropic和OpenAI员工实际相信的观点。
- 辞职还是不辞职。
- 安静地辞职是一种被支配的选择。
- 当你辞职时,非常严肃的人会明白这意味着什么。
- Jacob Coxon认为存在性风险很高,这就是他辞职的原因。
- Evan Hubinger认为存在性风险很高,这就是他留下的原因。
- Anthropic和OpenAI有商业动机去淡化存在性风险,而不是宣传它们。
- 我们现在该怎么做?
- 好吧,但AI究竟会如何杀死所有人?
- 最好相信科幻故事,因为你正身处其中。
- 彻底灭绝并不比失控难多少。
- 阴谋论小镇永远在招聘。
- 现在你看到了它。
Jacob Coxon Resigns From Anthropic In Protest And Sounds The Alarm
Jacob Coxon 辞职抗议 Anthropic 并发出警报
Jacob Coxon spent the last three years doing pretraining research at both OpenAI and Anthropic. He has come to realize that everyone involved is being wildly irresponsible.
Jacob Coxon 在过去三年里,先后在 OpenAI 和 Anthropic 从事预训练研究。他逐渐意识到,所有相关人员都极其不负责任。
He warns us: They are racing straight to superintelligence and gambling with our lives. I agree with and strongly endorse his statement.
他警告我们:他们正全速冲向超级智能,并将我们的生命置于赌局之中。我赞同并强烈支持他的声明。
If anything he sounds like an optimist. He’s asking you to consider what the next few years will actually feel like, which means he thinks you have a few years left.
如果说有什么乐观成分的话,那便是他的语气。他让你考虑未来几年实际会是什么感受,这意味着他认为你们还有几年的缓冲时间。
Jacob Coxon (former Anthropic and OpenAI, 160m+ views, September 8): I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
Jacob Coxon(前 Anthropic 和 OpenAI 员工,视频播放量超 1.6 亿,9 月 8 日):我今天从 Anthropic 辞职了。我在过去三年里,先后在 OpenAI 和 Anthropic 从事预训练研究。这两家公司都没有负责任地行事。它们正在全力冲刺自我改进的超级智能,并将我们的生命置于赌注之中。更多想法如下。
Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.
不要低估这项技术的力量。这些系统很快将成为超人级的存在,能够入侵任何事物、在一夜之间颠覆任何领域,并获取真正的权力和资源。我们都见证了这些领域的进步,而且进步并未放缓。
The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible – but I hear the same people express fear privately. No other human activity poses this level of danger.
认真构建 AI 的人们真心相信,到本十年末,AI 可能会让我们全部丧命。这不是营销噱头。事实上,许多高管和高级研究人员会在媒体面前用更审慎措辞来显得理智——但我听到这些人私下里表达了恐惧。没有其他人类活动能带来如此程度的危险。
A common response is “if they truly believe this, why are they still building it?” At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first – they believe no one else will act responsibly, so they must do it themselves, despite the risk.
一个常见的回应是:“如果他们真的这么认为,为什么还在继续构建?”在 OpenAI,许多人尚未深刻内化文明层面的利害关系。在 Anthropic,利害关系已被充分理解,但他们陷入了争先恐后的竞赛中——他们认为别人都不会负责任地行动,因此尽管有风险,他们必须亲自去做。
Accepting this race and entering the “endgame” is a hubristic gamble that should not be launched from a private company’s Slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available.
接受这场竞赛并进入“终局”是一种傲慢的赌博,不应从一家私营公司的 Slack 频道中发起。试图加速对齐过程,需要拥有非凡的信心,确信没有其他更好的路径可供选择。
I am optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities.
我对协调合作的可能性持乐观态度。像 Hugging Face 遭攻击这样的警示性事件,使得美国实验室之间的节奏协议变得更加可行。我感觉我们并没有走在防止全球竞赛的道路上,而这可能需要采取诸如暂时禁止提升模型能力等代价高昂的行动。
If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because “it’s happening anyway” – or take this moment to call for different conditions?
如果你是实验室的研究人员,我恳请你认真思考未来几年究竟会是什么感受。你是否想在缺乏对其心智的严谨理解的情况下,启动一个超级智能强化学习(RL)的运行?你是应该因为“反正事情已经在发生”而低头沉默——还是抓住这个时刻呼吁改变条件?
Jacob Coxon (WSJ interview): We’re on track for a lot of the most aggressive of these scenarios where by the end of next year things could be out of control already.
Jacob Coxon(华尔街日报采访):我们正朝着许多最激进的 scenarios 发展,到明年这个时候,局势可能已经失控。
If you want Jacob Coxon’s full views, I recommend his interview with Wired’s Maxwell Zeff. This thread has extensive quotes.
如果你想了解 Jacob Coxon 的完整观点,我推荐他接受《连线》杂志 Maxwell Zeff 的采访。该推文线程中包含大量引用。
Here is Jacob Coxon doing a 5 minute interview with Anderson Cooper. He speaks well and plainly, and it is clear how much the events of the last two months have made it much easier to speak plainly to a civilian like Cooper about what is happening.
这是 Jacob Coxon 接受 Anderson Cooper 5 分钟采访的视频。他表达清晰、直白,很明显过去两个月的事件使得向 Cooper 这样的平民更直白地谈论正在发生的事情变得容易得多。
Here is Jimmy Kimmel doing four minutes on this. He gets it. How is this not the top news story on every site, indeed.
这是 Jimmy Kimmel 对此进行的四分钟报道。他完全理解了这一点。这难道不是每个网站头条新闻吗?确实如此。
Yes, this is a common view, even if few have the courage to act.
是的,这是一种普遍的观点,尽管很少有人有勇气采取行动。
Alex Turner: I left Google DeepMind in June. Jacob is right: many researchers believe they are building something that could kill everyone on the planet. It was literally my day job to think about how to stop that.
Alex Turner:我于六月离开了 Google DeepMind。Jacob 说得对:许多研究人员相信他们正在构建某种可能杀死地球上所有人的东西。我的日常工作就是思考如何阻止这种情况发生。
That tells you how bad Alex Turner thought DeepMind’s actions were with regard to the Department of War. His day job was that he got paid by Google to think about how to stop AI from killing everyone, and he felt morally obligated to quit in protest.
这说明 Alex Turner 认为 DeepMind 在与战争部(Department of War)的关系中行为有多恶劣。他的日常工作是受雇于 Google,思考如何阻止 AI 杀死所有人,他觉得在道德上有义务辞职以示抗议。
Derek Thompson here writes about this as part of AI Safety Is Having a Moment.
Derek Thompson 在此撰文讨论此事,作为“AI 安全正处于风口浪尖”的一部分。
If you want to see the full list of lab employee quotes from the preference cascade, I compiled them into another post today.
如果你想查看偏好级联(preference cascade)中所有实验室员工的引语列表,我今天将它们整理到了另一篇文章中。
Mainstream Media Finally Pays Attention
主流媒体终于开始关注
There was strong coverage of this, starting at the Wall Street Journal: Anthropic Researcher Quits Over ‘Out-of-Control’ AI Fears. I do worry there is an unfortunate second way to read that headline.
对此进行了广泛报道,始于《华尔街日报》:Anthropic 研究员因担心 AI “失控”而辞职。我确实担心这个标题存在一种不幸的第二种解读方式。
Here is a sample of others, Astra can find many more:
以下是一些其他媒体的例子,Astra 可以找到更多:
CNN: ‘Gambling with our lives’: Another AI employee quits over safety concerns
CNN:“拿我们的生命赌博”:又一名 AI 员工因安全问题辞职
FT: Anthropic researcher quits over AI labs ‘gambling with our lives’
金融时报:Anthropic 研究员因 AI 实验室“拿我们的生命赌博”而辞职
Axios: Anthropic insiders warn AI could kill all humans.
Axios:Anthropic 内部人士警告 AI 可能杀死全人类。
Fortune: Anthropic researcher resigns, warning AI companies are ‘gambling with our lives’
财富杂志:Anthropic 研究员辞职,警告 AI 公司正在“拿我们的生命赌博”
Fox Business: Anthropic researcher says AI has over 10% chance to ‘kill all humans’
Fox Business:Anthropic 研究员称 AI 有超过 10% 的几率“杀死全人类”
BBC: Anthropic researcher believes more than 10% chance AI ‘could kill all humans.’
BBC:Anthropic 研究员认为 AI “可能杀死全人类”的概率超过 10%。
Preference Cascade at Anthropic
Anthropic 的偏好级联
Jacob Coxon started this. Evan Hubinger had the other key Tweet that set this off:
Jacob Coxon 发起了这一讨论。Evan Hubinger 发表了另一条关键推文,引发了这场风波:
Evan Hubinger (Alignment Science Lead, Anthropic): Jacob [Coxon] is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
Evan Hubinger(Anthropic 对齐科学主管):Jacob [Coxon] 说得对——我们确实真诚地认为 AI 可能会杀死全人类!我个人认为未来十年内发生的可能性超过 10%。我相信 Anthropic 正在全力以赴,但我们尚未制定解决超级智能对齐问题的计划,且目前看来并未明确走在正确的轨道上。
To be clear, as we say in our latest Risk Report, I think the risk from present models is low. What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought.
要明确的是,正如我们在最新《风险报告》中所言,我认为当前模型带来的风险很低。我担心的是由递归自我改进产生的超级智能,正如我们所说,其发展速度比我们预期的要快。
Indeed, 10% would be optimistic. Evan Hubinger said over 10%.
事实上,10% 已经是乐观估计了。Evan Hubinger 表示超过 10%。
One past statement of his estimates the chance of overall existential risk at 80%.
他过去曾有一项声明估计整体生存性风险的概率为 80%。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力