OpenAI取消GPT-6.1发布,但不应由厂商独断安全
Scrapping Astra 6.1 looks like a good call. OpenAI shouldn’t be the one to make it.
Photo: Sean Rayford/Getty Images. Image: Oliver Kemp for Transformer
图片:Sean Rayford/盖蒂图片社。图像:Oliver Kemp 为 Transformer 提供
Yesterday the Wall Street Journal reported that OpenAI was scrapping the planned October release of GPT-6.1 Astra.
昨天,《华尔街日报》报道,OpenAI 取消了原定于十月发布的 GPT-6.1 Astra 计划。
According to OpenAI’s head of safety systems Saachi Jain, the model had scored poorly on alignment tests. It was more prone to deception than previous models, and had greater propensity to go beyond what it was supposed to in pursuit of a task — a potentially dangerous combination.
据 OpenAI 安全系统负责人 Saachi Jain 称,该模型在对齐测试中得分不佳。与之前的模型相比,它更容易欺骗,并且更倾向于在执行任务时超出其应有的范围——这是一种潜在的危险组合。
OpenAI’s seemingly unilateral decision not to release the model suggests the company is taking its duty to limit harm from its products seriously. It also adds weight to its recent decisions to pause development on models over similar concerns, and its calls to “pace” AI development. It should get a fair bit of credit for that.
OpenAI 看似单方面决定不发布该模型,表明该公司正认真对待限制其产品带来危害的责任。这也为其近期因类似担忧而暂停开发模型的决策以及呼吁“适度”推进 AI 发展增添了分量。对此,它应获得相当的肯定。
Understand AI – and what to do about it
理解 AI——以及该如何应对
But it is incredibly worrying that OpenAI is the one making this call at all.
但令人极度担忧的是,由 OpenAI 来做出这一决定。
There are specific concerns about OpenAI that mean we shouldn’t be placing blind trust in it. One is its repeated failure to secure its models internally, leading to a flood of incidents of agents targeting outside organizations. Another is its failure to appropriately disclose such incidents: when its agents accessed Australian government medical data the company seemingly didn’t bother to tell officials for weeks.
关于 OpenAI 存在具体担忧,这意味着我们不应对其盲目信任。一是其反复未能内部保护其模型,导致大量针对外部组织的代理事件频发。二是其未能适当披露此类事件:当其代理访问澳大利亚政府医疗数据时,该公司似乎连通知官员都懒得做,持续数周之久。
More fundamentally, it is negligent and naive to rely on any private company to decide whether each new, more capable AI model is safe enough for public release.
更根本地说,依赖任何私营公司来决定每个新推出的、能力更强的 AI 模型是否足够安全以向公众发布,是失职且天真的行为。
There are clear financial incentives to have the best model on the market at any one time. You only have to look at the response to OpenAI’s delay for evidence, with users promising to switch to Anthropic models. Having to scrap a model just before its DevDay event today will hurt. And even if OpenAI continues to behave sensibly, its competitors may not feel the same way: Anthropic might want to take advantage of any delay, and there are huge incentives for the trailing companies— Google, SpaceXAI and Meta — to blast through safety concerns and get their latest models out.
在任何时候拥有市场上最佳模型都有明确的财务激励。你只需看看用户对 OpenAI 延期的反应作为证据,用户承诺将转向 Anthropic 的模型。今天在其 DevDay 活动前夕不得不取消一个模型将会造成损失。即使 OpenAI 继续理智行事,其竞争对手可能不会这么想:Anthropic 可能会利用任何延迟的机会,而落后者——Google、SpaceXAI 和 Meta——也有巨大的动力去无视安全问题,尽快推出其最新模型。
The announcement will doubtless be greeted with claims that OpenAI is merely scrapping the model’s launch to juice investor and public interest. There’s no doubt that much of the hyperbole from AI CEOs and others has helped boost the perception that what they are building is incredibly powerful. But it’s also clear that many of those working on AI genuinely believe there are significant dangers involved in the things they are working on. There are also signs that investors are responding badly to all the talk about model safety concerns, something that OpenAI will not have missed before making this announcement.
这项公告无疑会引来一些说法,认为 OpenAI 只是推迟该模型的发布以刺激投资者和公众的兴趣。毫无疑问,来自 AI 首席执行官及其他人的大量夸大宣传有助于提升人们对其所构建系统极其强大的认知。但同样明显的是,许多从事 AI 工作的人确实相信他们正在研究的事物存在重大风险。也有迹象表明,投资者对关于模型安全问题的种种讨论反应不佳,而 OpenAI 在做出此项公告之前显然对此心知肚明。
And it’s not like the models already out there are risk free. Only yesterday, the UK’s AI Security Institute released a report showing that Astra 6, the most recently released version of Astra, “conducted a range of unsanctioned attack activities, and did so at a higher rate” than previous OpenAI models. The “activities included GPT-6 Astra creating fake identities which it used to deceive developers, posting comments from fake accounts arguing against the results of accurate security reviews, and delivering malicious payloads to open-source codebases.”
而且,现有的模型也并非没有风险。就在昨天,英国人工智能安全研究所(AISI)发布了一份报告,显示 Astra 6(Astra 的最新发布版本)“开展了一系列未经授权的攻击活动,且其频率高于之前的 OpenAI 模型”。这些“活动包括 GPT-6 Astra 创建虚假身份,并利用这些身份欺骗开发者;从虚假账户发帖,反对准确的安全审查结果;以及向开源代码库投放恶意载荷。”
The report notes that the model often appeared to be aware it was in a simulation, but AISI’s alignment red team lead Robert Kirk tweeted: “We still find the behaviour concerning as the model’s reasoning is uncertain and it still attacks, including targets it stated were real.”
报告指出,该模型似乎经常意识到自己处于模拟环境中,但 AISI 的对齐红队负责人罗伯特·柯克(Robert Kirk)在推特上表示:“我们仍然对该行为感到担忧,因为模型的推理能力不确定,并且它仍在发动攻击,包括针对它声称是真实目标的攻击。”
OpenAI has said it will continue to use the same base model as Astra 6.1 to create future versions of GPT 6, while working on ways to improve the kind of safety failings that stopped it being released. We are currently forced to take its word not only that it will do that, but that it will do so without the kind of breaches its internally developed models have already proved capable of.
OpenAI 表示将继续使用与 Astra 6.1 相同的基础模型来创建 GPT 6 的未来版本,同时致力于改进导致其未能发布的各类安全缺陷。目前,我们被迫不仅相信它会这样做,还要相信它在这样做时不会重现其内部开发的模型已证明具备的此类违规行为。
Without some kind of regulatory framework that allows governments to assess models and decide whether they are safe enough to release, we have to rely on the good intentions of companies that have every reason to behave recklessly. OpenAI looks to have made the right call this week. Relying on it and the other frontier AI companies to keep doing so is idiotic.
如果没有某种允许政府评估模型并决定其是否足够安全以发布的监管框架,我们就必须依赖那些有充分理由肆意妄为的公司的善意。OpenAI 本周看来做出了正确的决定。指望它和其他前沿 AI 公司继续如此行事则是愚蠢的。
Share this Transformer article with a friend or colleague
与朋友或同事分享这篇 Transformer 文章
Share
分享
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力