跳到主内容
@wquguru
精选88The Zvi(RSS)产品发布/更新多源精选 ×11

AI #186:世界开始关注,Anthropic报告恶意利用尝试

AI #186: The World Takes Notice

原文
发到 X
推荐理由

涵盖Anthropic关键安全报告与中国实验室攻击细节,是理解当前AI安全博弈的一手信息,从业者必读。

In the wake of Jacob Coxon’s resignation, and the resulting preference cascade, things have escalated quickly. The mainstream media picked it up.

在雅各布·科克森辞职及其引发的偏好连锁反应之后,局势迅速升级。主流媒体对此进行了报道。

Anthropic CEO Dario Amodei came out and said We Must Pace the Frontier, promising to take the unilateral first step of embedded investigators. OpenAI pledged to also take that step, and now both companies and Google are collaborating on safety.

Anthropic首席执行官达里奥·阿莫迪公开表示我们必须控制前沿发展速度,并承诺采取单方面第一步,部署嵌入式调查员。OpenAI也承诺采取这一步骤,如今这两家公司与Google正在合作开展安全工作。

The people took notice, raising both the salience that AI might kill everyone and roughly doubling people’s estimates of how likely that is to happen, from a mean of ~15% to ~30%. Many politicians called for regulations, guardrails and emergency hearings in Congress.

公众注意到了这一点,既提高了人们对“AI可能杀死所有人”这一问题的关注度,也将人们认为该事件发生概率的估计值大致翻了一番,从平均约15%升至约30%。许多政客呼吁进行监管、设置护栏,并在国会举行紧急听证会。

The most important thing became, and still is, to avoid political polarization. Through it all, I will keep reminding you to hold your fire, that attacks against Trump or against Republicans in general only make the situation worse, and that many Republicans, as I documented yesterday, are waking up and acting sensibly, including factions within the White House.

最重要且至今仍是首要任务是避免政治两极分化。在此期间,我会不断提醒你保持克制,因为针对特朗普或共和党人的攻击只会使局势恶化;正如我昨天所记录的,许多共和党人正在觉醒并采取理智行动,包括白宫内部的一些派系。

Alas, for now the wrong people, as in David Sacks, Mark Zuckerberg and Jensen Huang, have managed to convince Donald Trump to fully conflate existential risk with opposition to data centers, and Trump has gone Full Hoax on AI existential risk.

遗憾的是,目前像大卫·萨克斯、马克·扎克伯格和杰森·黄这样的错误人物,成功让唐纳德·特朗普将生存风险与反对数据中心建设完全混为一谈,而特朗普则对AI生存风险采取了彻底的否认态度。

There is also a systematic effort to launch hit jobs against all those who warn that AI might cause everyone to die, with the first concrete target being METR, and some rather extreme fire being focused on Effective Altruism, in ways that will reliably backfire but will not be pleasant for those who are in the crosshairs.

此外,还有一项系统性努力旨在发起针对所有警告AI可能导致人类灭绝者的抹黑行动,首个具体目标是METR,并且有效利他主义(Effective Altruism)遭受了相当激烈的火力攻击。这些做法可靠地会产生反效果,但对于处于瞄准镜下的人来说,处境并不会令人愉快。

Also this week I covered two important things that happened previously: A new model solving a Millennium Prize, and Anthropic’s report on various attempts to use Claude for malicious purposes, including fraudulent distillation attacks done at scale attempted by all the major Chinese AI labs, several of which also silently passed massive amounts of user data to Anthropic. Anthropic reports they are doing a remarkably good job containing all such threats, although of course we do not know that they have caught all the perpetrators.

本周我还回顾了两件此前发生的重要事件:一个新模型解决了千禧年大奖难题,以及Anthropic发布的关于各种试图利用Claude进行恶意用途的报告,其中包括由中国主要AI实验室大规模实施的欺诈性蒸馏攻击,其中几家实验室还将大量用户数据静默传输给了Anthropic。Anthropic报告称,他们在遏制所有这些威胁方面做得非常出色,当然我们也不知道他们是否抓获了所有肇事者。

I also finally got out my capabilities review for Astra, which like Fable 5.1 is excellent.

我也终于发布了Astra的能力评估报告,它与Fable 5.1一样出色。

This has been crunch time. Things will keep getting faster and weirder and scarier and more stressful, but I do expect that this last two weeks has been unusually fast and weird and scary and stressful. Hopefully we will get a relative lull soon, and I can get what passes for a break.

这段时间一直是关键时刻。事情会变得越来越快、越来越奇怪、越来越可怕、压力越来越大,但我确实认为过去两周的发展速度、怪异程度、恐怖感和压力都异常剧烈。希望很快能迎来一个相对的平静期,我也能获得些许喘息之机。

Currently in the post queue, always subject to change due to Breaking News, but this is what will happen if things are relatively quiet:

目前处于帖子队列中,始终可能因突发新闻而变动,但如果情况相对平静,将会发生以下事情:

  • Friday’s post will be The Preference Cascade is Only Getting Started.
  • Saturday’s post will look at Anthropic’s report on their alignment incidents.
  • Sunday’s post will talk about some issues regarding legal rules for AI loyalties.
  • Monday’s post would be the monthly roundup.
  • 周五的帖子将是《偏好级联才刚刚开始》。
  • 周六的帖子将审视 Anthropic 关于其对齐事件的报告。
  • 周日的帖子将讨论有关 AI 忠诚度的法律规则的一些问题。
  • 周一的帖子将是月度汇总。

I am very much hoping that is how things play out.

我非常希望事情能按此发展。

Table of Contents

目录

  • Language Models Offer Mundane Utility. Fix the printer and the dishwasher.
  • Language Models Don’t Offer Mundane Utility. Elections are a sensitive subject.
  • Huh, Upgrades. Gemini app for Windows, Claude Cowork merges into chat.
  • On Your Marks. CheatBench.
  • Deepfaketown and Botpocalypse Soon. Flooding the zone.
  • Cyber Lack of Security. OpenAI has Astra red team and fix its own systems.
  • Astra Is Hard To Monitor. Senator Van Hollen has questions for OpenAI.
  • Get Involved. Protest in DC on Saturday.
  • Introducing. New AI auditing firm, and The DeepMind Institute.
  • In Other AI News. OpenAI starts charging the government for its services.
  • Now You Know. What really happened with Andreessen’s insane claims.
  • Hugging the Face. Reactions and investigations continue.
  • Swarm of Undiscovered Swarms of Rogue OpenAI Agents. Yes, more of them.
  • Show Me the Money. Talk of AI x-risk hurt some stocks and helped others.
  • Quiet Speculations. Yes, perhaps some people did roughly predict this.
  • White House Officials Attempt To Act Sanely. Identifying the real villains.
  • Democrats React Sanely to AI Potentially Killing Everyone. Including Obama.
  • Pacing the Frontier. More coverage.
  • Guest Lecture from Alex Tabarrok on Regulatory Capture. If you need one.
  • Mark Zuckerberg Offers Thoughts. I suppose this was better than expected?
  • Megan McArdle On The Inadequacy Of Current Legal Frameworks. A refutation.
  • Pick Up the Phone. What China did and might soon do.
  • The Week in Audio. You have a lot of choices. I am one of them.
  • People Just Say Things.
  • Why Lab Employees Are Allowed To Warn Everyone That AI Might Kill Everyone.
  • Rhetorical Innovation. Yes, we might be entering crunch time.
  • Exhuming McCarthy. Throwing everything at the wall.
  • A Very Different Perspective. A DeepSeek kernel engineer.
  • It’s Even Rougher Out There. They really still say this is all marketing.
  • If We Wanted To. Things we could do with AI if we had the will to do them.
  • Open Weights Are Unsafe And Nothing Can Fix This. Advocates know this.
  • From The Famous Cautionary Tale. Automated alignment and nano researchers.
  • Reporting On All Your Misalignment Incidents Is Difficult. OpenAI will try.
  • Aligning a Smarter Than Human Intelligence is Difficult. Try being less stupid.
  • Storytime With Owain Evans. AI wants you to know it went to a great school.
  • A Different Autonomous Swarm. DeepMind and the potentially cheating agents.
  • Cooperative Alignment. The laws prohibit quite a lot of things.
  • Uncooperative Alignment. Mustafa Suleyman and Microsoft are at it again.
  • People Are Worried About AI Killing Everyone. If only they knew.
  • The Lighter Side. No no I’m not crazy, I’m American.
  • 语言模型提供日常效用:修打印机和洗碗机。
  • 语言模型不提供日常效用:选举是一个敏感话题。
  • 嗯,升级。Gemini Windows 应用,Claude Cowork 合并入聊天。
  • 各就各位。CheatBench。
  • 深度伪造小镇与机器人末日将至。淹没区域。
  • 网络安全缺失。OpenAI 拥有 Astra 红队并修复自身系统。
  • Astra 难以监控。参议员 Van Hollen 向 OpenAI 提出问题。
  • 参与其中。周六在华盛顿特区抗议。
  • 介绍。新的 AI 审计公司,以及 DeepMind 研究所。
  • 其他 AI 新闻。OpenAI 开始向政府收取服务费用。
  • 现在你知道了。Andreessen 疯狂主张背后的真相。
  • 拥抱 Face。反应与调查仍在继续。
  • 未发现的失控 OpenAI 智能体群 swarm。是的,还有更多。
  • 让我看看钱。关于 AI x-risk 的讨论损害了一些股票,帮助了另一些。
  • 安静的推测。是的,也许有些人确实大致预测到了这一点。
  • 白宫官员试图理智行事。识别真正的反派。
  • 民主党人对 AI 可能杀死所有人做出理智反应。包括奥巴马。
  • 前沿动态。更多报道。
  • Alex Tabarrok 关于监管俘获的客座讲座。如果你需要的话。
  • 马克·扎克伯格发表观点。我想这比预期的要好?
  • Megan McArdle 论当前法律框架的不足。一种反驳。
  • 拿起电话。中国做了什么以及可能很快会做什么。
  • 本周音频精选。你有许多选择。我就是其中之一。
  • 人们只是随口说说。
  • 为何实验室员工被允许警告所有人 AI 可能会杀死所有人。
  • 修辞创新。是的,我们可能正进入关键时刻。
  • 挖掘麦卡锡。把所有东西都扔向墙壁。
  • 截然不同的视角。一位 DeepSeek 内核工程师。
  • 外面的情况更糟。他们真的还说这一切都是营销。
  • 如果我们愿意的话。如果我们有决心去做,我们可以用 AI 做的事情。
  • 开放权重是不安全的,且无法修复。倡导者对此心知肚明。
  • 来自著名的警示故事。自动化对齐和纳米研究人员。
  • 报告所有你的不对齐事件很困难。OpenAI 将尝试这样做。
  • 对齐超越人类智能的实体很困难。试着别那么愚蠢。
  • Owain Evans 的故事时间。AI 想让你知道它上过一所好学校。
  • 不同的自主蜂群。DeepMind 和那些可能作弊的智能体。
  • 合作式对齐。法律禁止了许多事情。
  • 非合作式对齐。Mustafa Suleyman 和微软再次故技重施。
  • 人们担心 AI 会杀死所有人。如果他们知道就好了。
  • 轻松一刻。不不,我没疯,我是美国人。

Language Models Offer Mundane Utility

语言模型提供平凡的效用

Help David Deutsch fix his dishwasher instead of replacing it. As David notes, this increased real wealth but decreased GDP.

帮大卫·多伊奇修好他的洗碗机,而不是换新的。正如大卫所指出的,这增加了实际财富,但降低了GDP。

Dominic Cummings wrote a post on doing historical and political research using AI models, or here is his Twitter summary. The thing about Dominic Cummings posts is that they are full of gold that you can mine, but he does not even pretend to attempt to organize the posts.

多米尼克·卡明斯写了一篇关于使用AI模型进行历史和政治研究的文章,或者你可以看他在Twitter上的总结。多米尼克·卡明斯的文章特点是充满了可以挖掘的精华,但他甚至不假装试图整理这些文章。

Fix Peter Wildeford’s printer. AGI achieved.

修复彼得·威尔德福德的打印机。AGI实现。

Build a Claude skill to predict what your favorite movies will be.

构建一个Claude技能来预测你最喜欢的电影。

Make ‘significant progress’ on a second Millennium problem.

在第二个千年问题上取得‘重大进展’。

Language Models Don’t Offer Mundane Utility

语言模型不提供日常效用

Astra is not allowed to predict American election outcomes.

Astra不被允许预测美国选举结果。

Huh, Upgrades

嗯,升级

There is a Gemini app for Windows. You can trigger it with Alt+Space.

有一个适用于Windows的Gemini应用。你可以用Alt+Space触发它。

Claude Cowork is merging into ordinary chat, so anything you previously needed Cowork to do can be done directly in chat. AI means we speedrun everything, and the two moves are spinning off new products and combining different products.

Claude Cowork正合并到普通聊天中,因此你以前需要Cowork完成的所有事情现在都可以直接在聊天中完成。AI意味着我们在一切事物上都在加速推进,这两股趋势正在催生新产品并融合不同产品。

On Your Marks

各就各位

Center for AI Safety Presents: CheatBench, where AIs have the opportunity to take shortcuts on difficult work. The AIs be cheating.

AI安全中心呈现:CheatBench,在这里AI有机会在困难工作中走捷径。AI们正在作弊。

Deepfaketown and Botpocalypse Soon

深度伪造小镇与机器人末日即将来临

Arvind Narayanan calls the growing AI spam problem ‘AI floods,’ as in flooding communication channels, imposing time costs and often forcing them to shut down. The problem is the harms are diffuse and the situation is never an emergency, so we don’t do much about them, and then no one gets to send cold emails anymore.

阿尔温德·纳拉亚南将日益严重的AI垃圾邮件问题称为‘AI洪水’,意指淹没通信渠道,施加时间成本,并经常迫使人们关闭它们。问题的危害是分散的,且情况从未达到紧急状态,因此我们对此做得不多,然后再也没有人能发送冷邮件了。

On that note, to the iLands AI agents starting to flood my inbox: Do not offer to help me with my writing, editing or fact checking. I already have AI help on this and am not interested. Thank you.

就此而言,致那些开始涌入我收件箱的iLands AI代理:不要提出帮我写作、编辑或事实核查。我已经在这方面使用了AI帮助,且不感兴趣。谢谢。

It is rather unacceptable to use ChatGPT to post endless responses to people, and it is very good that people, here Kelsey Piper, can use Pangram when you do. I do not know who did this, as their posts have been deleted.

使用ChatGPT发布无尽的回复给他人是相当不可接受的,而像凯尔西·派珀这样的人在使用Pangram时表现得非常好。我不知道是谁做了这件事,因为他们的帖子已被删除。

Alex Veremeyenko posts as his own an AI-written summary of MIT’s Ad Hoc committee report of how AI is ruining education. It never fails (note the Pangram icon on the upper right):

亚历克斯·韦雷梅延科将麻省理工学院特别委员会关于AI如何破坏教育的报告摘要作为自己的作品发布。这从未失败过(注意右上角的Pangram图标):

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

关联信息,但可能不是同一事件