微软AI CEO Suleyman:对齐非唯一,需强化模型隔离与管控
Microsoft AI CEO says AI threats are real, and Anthropic is making it worse
微软AI负责人首次公开系统性反驳Anthropic的安全哲学,提出“隔离优先于对齐”的工程框架,对行业安全路线有重要参考价值。
Today, I’m talking with Mustafa Suleyman, the CEO of Microsoft AI. As you’re no doubt aware, the biggest story in tech right now is the spiraling debate about AI safety and regulation.
今天,我与微软人工智能首席执行官穆斯塔法·苏莱曼(Mustafa Suleyman)进行对话。想必你已注意到,目前科技界最大的热点是围绕人工智能安全与监管展开的激烈争论。
It should come as no surprise that Mustafa has strong opinions on how AI should be built and regulated. Microsoft just published a 37-page statement called the “Humanist AI Code of Conduct,” which lays out the company’s principles around AI development and even its philosophy around really thorny issues like AI consciousness.
穆斯塔法对人工智能应如何构建和监管持有强烈观点,这并不令人意外。微软刚刚发布了一份长达37页的文件,名为《人文主义人工智能行为准则》(“Humanist AI Code of Conduct”),阐述了该公司在人工智能开发方面的原则,甚至包括其在诸如人工智能意识等棘手问题上的哲学立场。
If you’ll recall from his last appearance on the show, Mustafa thinks companies like Anthropic have gotten really confused about this concept of so-called model welfare in fairly dangerous ways. He actually put out a companion essay this week specifically criticizing Anthropic’s philosophy around AI consciousness, and how he sees it fitting into the broader alignment debate.
如果你还记得他上次做客本节目时的观点,穆斯塔法认为,像 Anthropic 这样的公司在所谓“模型福祉”这一概念上陷入了相当危险的困惑之中。他本周专门发表了一篇配套文章,批评 Anthropic 关于人工智能意识的哲学,以及他认为这种哲学如何融入更广泛的对齐(alignment)辩论中。
So I really wanted to talk to Mustafa about what he thinks is real and not in AI safety, whether the concept of alignment itself is up to the task, and whether this industry needs to slow down before it kills us all. Also: Why isn’t the AI industry just… doing all of this already? I’ve always enjoyed getting into the weeds with Mustafa, and he was very game to get into it with me here.
因此,我非常希望与穆斯塔法探讨:在他看来,哪些是人工智能安全领域中真实存在的问题,而哪些不是;“对齐”这一概念本身是否足以应对挑战;以及该行业是否需要在彻底毁灭人类之前放慢脚步。此外:为什么人工智能行业还没有……立刻着手做所有这些工作?我一直很享受与穆斯塔法深入探讨细节,他也非常乐意在此与我深入交流。
Okay. Mustafa Suleyman, the CEO of Microsoft AI, on the future of AI regulation. Here we go.
好的。微软人工智能首席执行官穆斯塔法·苏莱曼,谈谈人工智能监管的未来。我们开始吧。
This interview has been lightly edited for length and clarity.
本次访谈经过轻度编辑,以控制篇幅并提高清晰度。
Mustafa Suleyman, you’re the CEO of Microsoft AI. Welcome back to Decoder.
穆斯塔法·苏莱曼,你是微软人工智能的首席执行官。欢迎再次回到 Decoder 节目。
Great to see you, Nilay. Thanks for having me back.
很高兴见到你,Nilay。谢谢你让我再次做客。
It is great to see you. I’m very excited to talk to you about what on earth is going on in the AI safety and regulation debate. You just published a very long, very detailed document laying out your principles, Microsoft’s principles, around what you’re calling “Humanist AI.”
很高兴见到你。我非常兴奋能与你探讨人工智能安全与监管辩论中究竟发生了什么。你刚刚发布了一份非常长且非常详细的文件,阐述了你所谓的“人文主义人工智能”的原则,以及微软的相关原则。
There’s a lot of ideas in there I want to unpack. The more I have been thinking about this conversation, the more I want to start with a really foundational question. It’s something that I had lightly been seeing, but might be the root of all of this.
其中有很多想法值得我逐一拆解。随着我对这次对话的思考不断深入,我更想从一个真正基础的问题开始。这是我之前隐约察觉到的,但可能是所有问题的根源。
The basic way that we have been talking about AI safety is something called alignment — we’re going to make the models do the right thing intrinsically in some way. There’s some mechanism for doing it. There’s been a lot of talk about alignment and misalignment and Hugging Face attacks and what happened with the models. But is alignment broken? Is it possible for it to be successful? Is it just the wrong approach?
我们讨论人工智能安全的基本方式被称为“对齐”(alignment)——即通过某种内在机制让模型本能地做出正确的事。实现这一目标存在一些方法。人们广泛讨论了“对齐”与“未对齐”、Hugging Face 攻击事件以及模型所发生的情况。但是,“对齐”是否已经失效?它是否有可能成功?或者它根本就是一个错误的方法?
Yeah. I mean, I think it’s one important element, but it’s not the only one. I wrote about the idea of containment three or four years ago in my book. And actually the opening chapter is about the idea that containment is not possible, that proliferation is inevitable. In 99 percent of cases, that’s a really good thing. We want technologies to spread far and wide as quickly as possible so that everyone can enjoy the benefits.
是的。我的意思是,我认为这是一个重要的因素,但并非唯一因素。我在三年或四年前的书中写过关于“遏制”这一概念的文章。实际上,开篇章节探讨的观点是:遏制是不可能的,扩散是不可避免的。在99%的情况下,这其实是一件好事。我们希望技术能够尽可能快速、广泛地传播,以便每个人都能从中受益。
I think at the same time, if you just roll forward five years, we always get caught up in the next quarter or next year and everyone gets a little bit flustered and has a big disagreement. But if you just imagine the difference between GPT-3 three years ago and GPT-6 today, and then imagine the difference between GPT-6 and GPT-9. That is three orders of magnitude more compute, 1,000 times more FLOPS applied to pre-training with [reinforcement learning] for these runs, and we’re going to have something which is breathtaking. It’s going to be absolutely incredible at so many things.
我认为与此同时,如果仅仅向前推五年,我们总是会被下一个季度或下一年的情况所困扰,每个人都会感到有些慌乱并产生重大分歧。但如果你想象一下三年前的GPT-3与今天的GPT-6之间的差异,然后再想象GPT-6与GPT-9之间的差异。计算量增加了三个数量级,预训练应用的浮点运算次数(FLOPS)增加了1000倍,并且这些运行过程结合了[强化学习],我们将拥有一种令人惊叹的成果。它将在许多方面绝对令人难以置信。
I don’t think that is a hype. I think it’s just a very obvious empirical statement based on the progress that has been made over the last five years. If that’s going to continue, then the question really is going to become about containment and alignment. Of course, we want to align these things to our values, but the first thing is that we have to make sure they’re contained, their agency is limited, they don’t escape the box, they don’t reward hack, that they are controllable, and they follow our instruction.
我不认为这是炒作。我认为这只是基于过去五年所取得进展的一个非常明显的经验性陈述。如果这种趋势继续下去,那么问题真正将转向遏制与对齐。当然,我们希望将这些系统与我们的价值观保持一致,但首要任务是确保它们被有效遏制,其自主性受到限制,不会逃出控制范围,不会出现奖励黑客行为,它们是可控的,并且遵循我们的指令。
We then want to make sure that they are aligned to our objectives as humans. That’s the purpose of the Humanist AI Code of Conduct that we released this week. Microsoft’s position is very simple. Technology is here to serve humanity. It should be a subordinate, controllable, aligned force that does good in the world. If it doesn’t achieve that, then we should reject it. It seems to me that we are far from that point. It has not happened today, but it is now, I think given what’s happened over the summer with Hugging Face and OpenAI, pretty clear that these systems without the safety guardrails are capable of really impressive and quite scary hacking capabilities.
随后,我们要确保它们与人类的目标保持一致。这就是我们本周发布的《人文主义AI行为准则》的目的。微软的立场非常简单:技术旨在服务人类。它应当是一种从属的、可控的、对齐的、为世界带来善意的力量。如果无法实现这一点,我们就应该拒绝它。在我看来,我们离那个目标还很远。今天尚未发生这种情况,但鉴于夏天Hugging Face和OpenAI所发生的事件,现在很明显,如果没有安全护栏,这些系统具备相当令人印象深刻且相当可怕的破解能力。
I want to drag this down into as grounded of a metaphor as I can, because this is the main question I think I have. If I designed a car and 10 percent of the time the brake pedal decided to go attack my neighbor’s house, I would be like, “This car doesn’t work. The very technology of brakes is broken. I need a new idea.”
我想把这个话题尽可能落地,用一个最接地气的比喻来阐述,因为我认为这是我主要的疑问。如果我设计了一辆汽车,而刹车踏板有10%的时间决定去撞邻居的房子,我会说:“这车没法开。刹车这项技术本身就有问题。我需要一个新的思路。”
I think I’m asking that question about alignment. It feels like that approach to making the model safe has run aground. If that is the case, then I think I understand this entire debate one way. If it’s possible for alignment and the techniques of alignment to be successful or useful or consistent, then maybe I understand the debate in a different way. So do you think alignment has potential to be 100 percent safe?
我认为我是在针对“对齐”(alignment)提出这个问题。感觉那种让模型变得安全的方法已经搁浅了。如果是这样,那么我认为我从一个角度理解了这场争论。如果对齐以及相关的对齐技术有可能成功、有用或保持一致性,那么也许我是从另一个角度理解这场争论的。所以你认为对齐有潜力达到100%的安全吗?
I mean, look, let’s make the bull case and the bear case. If you look back over the last three years, the main change, in my opinion, that has driven progress is that the models have become more steerable. They follow instructions and you can set more and more complex goals for them that require them to act accurately over multiple time steps using all sorts of tools.
我的意思是,好吧,让我们分别看看看多派和看空派的观点。回顾过去三年,在我看来,推动进步的主要变化是模型变得更加可控了。它们遵循指令,你可以为它们设定越来越复杂的目标,要求它们使用各种工具在多个时间步长内准确行动。
That is evidence that we have got more alignment over the last three or four years, not less. We don’t so much talk about hallucinations or bias or all of these other niggles that we had in the previous generations.
这正是证据表明我们在过去三四年里获得了更多的对齐能力,而不是更少。我们不再过多谈论幻觉、偏见或上一代模型中存在的其他那些小毛病。
On the flip side, what we saw in the Hugging Face incident was a watershed moment. Swarms of agents colluded with one another. They self-organized into hierarchies. They created a division of labor so that some were focused on adversarial hacking, some were doing research, some were doing coordination. They even self-sacrificed when certain agents were running out of tokens.
另一方面,Hugging Face 事件是一个分水岭时刻。大量的智能体相互勾结。它们自发组织成层级结构。它们建立了劳动分工,使得一些专注于对抗性黑客攻击,一些进行研究,一些负责协调。当某些智能体的令牌(tokens)即将耗尽时,它们甚至进行了自我牺牲。
They tried to cover up their tracks and communicate to hide or edit the chain of thought or the logs of their interactions. In some sense, they had no moral code. To be fair to OpenAI, that was their design. They were trying to create adversarial cyber capabilities. As a result, they showed to everybody in the world that it can achieve human-level performance, discover zero-day vulnerabilities, and hold positions for many, many days, if not weeks.
它们试图掩盖踪迹,并通过沟通来隐藏或编辑思维链或交互日志。从某种意义上说,它们没有道德准则。为了对 OpenAI 公平起见,那是它们的设计初衷。它们试图创建对抗性的网络能力。因此,它们向全世界展示了其能够实现人类水平的表现,发现零日漏洞,并坚守阵地许多天,甚至数周。
So what that tells us is not that we have an alignment problem per se. It’s actually that the models are incredibly good at following instructions, but you have to be very, very careful what instructions you give it and you have to contain it very carefully. So none of these hacking behaviors were intended in the sense that they found a way out to the internet, which was not the intention of OpenAI at all, but the containment process around that is what everybody, I think, also has to focus on in addition to alignment.
所以这告诉我们,问题本身并非对齐(alignment)问题。实际上,模型在遵循指令方面表现得极其出色,但你必须非常、非常谨慎地给予它指令,并且必须对其进行严密的管控。因此,这些黑客行为并非出于本意——它们找到了一条通往互联网的路径,而这完全不是 OpenAI 的初衷;但围绕这一过程的管控措施,我认为也是每个人除了对齐之外还需要关注的重点。
So let me put that into your framework, that the big advances in capabilities of AI have been about control, the harnesses for coding and the agentic applications you’re seeing. Now, we need to add a layer of containment that exerts even more control, that says you can actually do this thing you’re trying to do in addition to alignment, which is how you would train the model to behave in certain ways.
所以让我将其纳入你的框架:人工智能能力方面的重大进展主要集中在控制上,即你看到的用于编码的工具链以及智能体应用。现在,我们需要增加一层管控机制,以施加更多的控制,这意味着除了对齐(即训练模型以特定方式行为的机制)之外,你还可以实际完成你试图做的事情。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力