AI周报185:Claude Fable 5.1与GPT-6
AI #185: Preference Cascade
深度解读了Claude Fable 5.1和GPT-6 Astra的核心差异与监控风险,特别是关于Astra可监控性下降的分析,对Agent开发者极具参考价值。
The world of AI is inside my OODA loop. Even if I can process all the incoming information and sculpt it into posts, and even using Saturday and Sunday as flex slots, I don’t have enough days of the week to post all the posts that need posting.
AI的世界已融入我的OODA循环。即便我能处理所有涌入的信息并将其塑造成帖子,甚至将周六和周日作为弹性时段,我在一周内也没有足够的天数来发布所有需要发布的帖子。
That was already true. There was already a preference cascade happening where people finally were admitting that they thought AI might well kill everyone.
这早已是事实。当时已经出现了一种偏好级联现象,人们终于承认他们认为AI可能会杀死所有人。
Then Jacob Coxon resigned from Anthropic, rang the warning bells and turned that cascade into an avalanche.
随后,Jacob Coxon从Anthropic辞职,拉响了警示铃,并将这种级联演变成了一场雪崩。
Now that is what everyone is talking about. Finally, everyone is actually saying the thing, out loud. I plan to cover that in its own post soon.
现在这是每个人都在谈论的话题。终于,每个人都在大声说出这件事。我计划很快在另一篇帖子中对此进行覆盖。
There are several things in the weekly that, in a normal week, would get their own coverage. Senator Sanders and Representative Casar introduced an outright ban on superintelligence and I have to remind myself that happened this week. Suddenly it is not so crazy to think such a thing might pass.
本周刊中有几件事,在正常的一周里会获得单独报道。参议员桑德斯和众议员卡萨尔提出了一项对超级智能的全面禁令,我必须提醒自己这确实发生在本周。突然之间,想到这样的事情可能通过就不再那么疯狂了。
So here’s what I’ve already posted about so far since the last weekly:
所以,自上一期周刊以来,我已经发布了以下内容:
- Claude Fable and Mythos 5.1: The System Card.
- Claude Fable and Mythos 5.1: Capabilities.
- OpenAI and the Wiki Incident.
- Reports of OpenAI agents compromising additional Wikis continue to stream in, and OpenAI is continuing not to be the one to disclose them.
- An Alien Mind: Jakub Pachocki Warns Us.
- Astra is Hard to Monitor.
- GPT-6 Astra: The System Card, Alignment and What Comes Next.
- Claude Fable与Mythos 5.1:系统卡片。
- Claude Fable与Mythos 5.1:能力。
- OpenAI与Wiki事件。
- 关于OpenAI代理破坏更多Wiki的报告仍在不断传来,而OpenAI继续不主动披露这些信息。
- 异星思维:Jakub Pachocki向我们发出警告。
- Astra难以监控。
- GPT-6 Astra:系统卡片、对齐以及后续发展。
Anthropic’s Claude Fable 5.1 is a very good model. OpenAI’s GPT-6 Astra is also a very good model, and a bigger improvement. Try both, see what works for you where.
Anthropic的Claude Fable 5.1是一个非常优秀的模型。OpenAI的GPT-6 Astra也是一个非常优秀的模型,且改进幅度更大。两者都试试,看看哪个更适合你。
The HuggingFace OpenAI saga got a prequel, as it turns out that there was a ‘Wiki Incident’ prior to the hack, which OpenAI decided not to disclose, that in some key ways changes our interpretation of the timeline.
HuggingFace上的OpenAI saga有了前传,事实证明,在黑客攻击之前就已经存在一个‘Wiki事件’,而OpenAI决定不予披露,这在某些关键方面改变了我们对时间线的解读。
OpenAI Chief Scientist Jakub Pachocki warned us that capabilities are developing rapidly, alignment is not keeping pace and monitorability is eroding fast. We will need to find ways to cooperate and slow down, or we are all cooked. This was a very good essay.
OpenAI首席科学家Jakub Pachocki警告我们,能力正在迅速发展,对齐进度未能跟上,可监控性正在迅速恶化。我们需要找到合作并放缓速度的方法,否则我们都完了。这是一篇非常好的文章。
The Astra system card and several other OpenAI sources had lots of important and useful information. But the Astra release announcement, and some parts of the system card and Jakub’s essay, treated the evidence as being far less alarming than it is, and overclaiming on alignment concerns in ways that make me worried they don’t understand the dangers, both on alignment and monitorability.
Astra的系统卡片以及其他几个OpenAI来源包含大量重要且有用的信息。但是,Astra的发布公告、系统卡片的部分内容以及Jakub的文章,将证据描述得远不如实际情况那样令人担忧,并在对齐问题上过度夸大,这让我担心他们并不理解在对齐和可监控性方面的危险。
Astra remains monitorable, contrary to some false alarms. but it is substantially less monitorable than Sol, in ways I do not think can be explained only by the increase in its capabilities. Right now OpenAI is extremely dependent on CoT monitoring, everyone else depends on it quite a lot as well, and it looks like it may not last much longer, and that Astra already is on the edge of steganographic capabilities and can do substantial obfuscation if and only if it thinks you would think it is up to no good. And Astra will ask the question previous misaligned AIs wouldn’t, as in it will not cheat in situations where it would expect to be caught.
与一些误报相反,Astra 仍然可被监控。但它在诸多方面远比 Sol 难以监控,我认为这些差异不能仅用其能力提升来解释。目前 OpenAI 极度依赖思维链(CoT)监控,其他机构也高度依赖该方法;看起来这种方法可能不会持续太久,而 Astra 已经处于隐写术能力的边缘,只要它认为你会觉得它图谋不轨,就能进行大量的混淆操作。并且 Astra 会提出之前未对齐的 AI 不会提出的问题,例如在预期会被抓到的情况下,它不会作弊。
Astra is substantially more aligned than Sol in the sense of what I call ‘mundane alignment,’ or the practical day to day use of the model. One might also call it ‘prosaic alignment.’ That is very different from alignment that scales, or that matters at the highest stakes, the one that ultimately counts, sometimes called ‘super alignment.’
就我所谓的‘日常对齐’(mundane alignment),即模型在日常实用中的表现而言,Astra 比 Sol 对齐程度高得多。也可以称之为‘平凡对齐’(prosaic alignment)。这与可扩展的对齐或在最高风险情境下才重要的对齐截然不同,后者才是最终关键的、有时被称为‘超级对齐’(super alignment)的东西。
That still leaves the pending standard post covering Astra’s Capabilities, as well as coverage of the Millennium Prize where an AI solved Navier-Stokes less then two weeks after OpenAI started training it, and we saw Anthropic and OpenAI unable to get along even there.
这仍留下了待发布的标准文章,涵盖 Astra 的能力,以及对千禧年大奖难题的报道:某 AI 在 OpenAI 开始训练它不到两周后便解决了纳维-斯托克斯方程,而我们看到 Anthropic 和 OpenAI 即便在此事上也无法和睦相处。
I also have not had opportunity to examine Anthropic’s alignment assessment of recent cybersecurity incidents, or to offer my thoughts around Alex Mallen’s excellent post.
我也尚未有机会审查 Anthropic 对近期网络安全事件的对齐评估,或针对 Alex Mallen 的优秀帖子发表我的看法。
Tentative schedule:
暂定日程:
- Friday: Coxon and the Preference Cascade.
- Saturday: The Millennium Prize and the new OpenAI model being trained.
- Sunday: Astra Capabilities.
- 周五:Coxon 与偏好级联。
- 周六:千禧年大奖难题与新训练的 OpenAI 模型。
- 周日:Astra 能力。
After that, we will see where we are.
之后,我们再看情况如何。
Table of Contents
目录
- Language Models Offer Mundane Utility. The research speeds up.
- Language Models Don’t Offer Mundane Utility. Limits of intelligence.
- Huh, Upgrades. DeepSeek 4.1-Flash exists.
- How To Tell a Fable. Some tips for optimizing Fable 5.1.
- On Your Marks. Fable 5.1 regresses on Eliezer’s fickle FictionPlotBench.
- Deepfaketown and Botpocalypse Soon. If they oppose Pangram, one guess why.
- Levels of Friction. The slow-moving crisis in mispriced restaurant reservations.
- Cyber Lack of Security. An example of something that could have happened.
- A Young Lady’s Illustrated Primer. If you can’t ban ‘em, join ‘em.
- They Took Our Jobs. Relational societies and other hells.
- Anthropic Offers Economic Scenarios. Playing with toy economic models is fun.
- Get Involved. Coefficient Giving wants to give out a lot of money. Like, a lot.
- Introducing. Apollo Research offers Watcher Live to guard your agents.
- In Other AI News. The HuggingFace price is right.
- Show Me the Money. Mistral raises $3 billion. In Europe that’s a lot.
- Quiet Speculations. The future is going to be absurd, even while living in it.
- The Quest for Sane Regulations. Calls for action prior to the preference cascade.
- The OpenAI Policy and Lobbying Department. Hope for a new leaf. We’ll see.
- Greetings From the Department of War. The government agents go a bit rogue.
- Hugging The Face. Coverage took some time, but we saw some good pieces.
- Hugging the Question. When Congress asks, you are supposed to answer.
- The Ban Artificial Superintelligence Act. Sanders and Casar are go all the way.
- Chip City. How Chinese firms evade US export controls.
- The Week in Audio. Roose, Marcus and yours truly.
- People Just Say Things.
- PauseAI Global Disendorsed PauseAI US. An unfortunate situation.
- Paul Christiano Joins Board of OpenAI Foundation. An excellent pick. The best.
- Rhetorical Innovation. We always knew these discussions would happen.
- Aligning a Smarter Than Human Intelligence is Difficult. Think like an Astra.
- Cooperative Alignment. Personas are not merely selected.
- Drive to Survive. Agents that have to earn their compute.
- People Are Worried About AI Killing Everyone. The ones before Coxon.
- Other People Are Not As Worried About AI Killing Everyone. We are toast.
- The Lighter Side. Oh.
- 语言模型提供日常效用。研究加速推进。
- 语言模型不提供日常效用。智能的局限。
- 嘿,升级了。DeepSeek 4.1-Flash 存在。
- 如何讲述寓言。优化 Fable 5.1 的一些技巧。
- 各就各位。Fable 5.1 在 Eliezer 善变的 FictionPlotBench 上出现倒退。
- 深度伪造镇与机器人末日将至。如果它们反对 Pangram,猜一个原因。
- 摩擦层级。餐厅预订定价缓慢演变的危机。
- 网络安全缺失。一个本可能发生的事例。
- 一位年轻女士的图解手册。如果不能禁止它们,那就加入它们。
- 它们夺走了我们的工作。关系型社会及其他地狱。
- Anthropic 提供经济场景。玩弄玩具经济模型很有趣。
- 参与其中。Coefficient Giving 想要发放大量资金。是的,非常多。
- 介绍。Apollo Research 提供 Watcher Live 以守护你的智能体。
- 其他 AI 新闻。HuggingFace 的定价很合理。
- 金钱至上。Mistral 融资 30 亿美元。在欧洲,这是一笔巨款。
- 静默推测。未来将变得荒诞不经,即便我们身处其中。
- 寻求理性监管。在偏好级联发生之前的行动呼吁。
- OpenAI 的政策与游说部门。期待翻开新篇章。拭目以待。
- 来自战争部的问候。政府智能体开始有些失控。
- 拥抱 HuggingFace。报道花了些时间,但我们看到了一些不错的文章。
- 拥抱问题。当国会提问时,你理应作答。
- 《禁止人工超智能法案》。桑德斯和卡萨尔等人主张彻底禁止。
- 芯片之城。中国企业如何规避美国出口管制。
- 本周音频内容。Roose、Marcus 和我本人。
- 人们随口乱说。
- PauseAI 全球版被否决。PauseAI 美国版获通过。情况令人遗憾。
- Paul Christiano 加入 OpenAI 基金会董事会。一个极佳的选择。最好的选择。
- 修辞创新。我们一直都知道这些讨论会发生。
- 对齐超越人类水平的智能是困难的。像 Astra 一样思考。
- 合作式对齐。角色设定并非仅仅被选中。
- 生存驱动。必须赚取计算资源的智能体。
- 人们担心 AI 会杀死所有人。Coxon 之前的那些观点。
- 其他人并不那么担心 AI 会杀死所有人。我们要完蛋了。
- 轻松一刻。哦。
Language Models Offer Mundane Utility
语言模型提供日常效用
Remember when people still tried to doubt that AI coding massively sped people up? Ruben Bloom, who was part of the METR uplift study that found devs were not much accelerated by AI, now reports he’s seeing unambiguous 10-50x speedups on projects. This was before Astra.
还记得人们曾试图质疑 AI 编程是否真的大幅提升了效率吗?Ruben Bloom 曾是 METR 提升研究的一部分,该研究发现开发者并未因 AI 获得显著加速;如今他报告称,在项目上看到了明确无误的 10-50 倍速度提升。这发生在 Astra 之前。
Claude submits a formal Lean proof of Fermat’s Last Theorem.
Claude 提交了一份费马大定理的正式 Lean 形式化证明。
Nathan : Can we have a speedrunning-style thing where the game is now to prove theorems in fewer characters of lean?
Nathan:我们能不能搞一个类似速通(speedrunning)的活动,目标是用更少的 Lean 字符数来证明定理?
Jakub Pachocki incidentally said in An Alien Mind that OpenAI could make the models better at math, but is choosing not to focus on that. Which means that the math progress we see is well short of what we could be seeing.
Jakub Pachocki 在《外星思维》(An Alien Mind)中顺便提到,OpenAI 本可以让模型在数学方面表现更好,但选择不专注于此。这意味着我们目前看到的数学进展远不及我们本可以达到的水平。
Language Models Don’t Offer Mundane Utility
语言模型并不提供平庸的效用
A zen koan: Are you sure that isn’t the problem, sir?
一则禅宗公案:先生,您确定那不是问题所在吗?
kache (reviewing Astra): I am shocked how little of my personal progress in everything at life was blocked by intelligence.
kache(评论 Astra):令我震惊的是,我在生活中各方面的个人进步,受智力因素阻碍的程度竟然如此之小。
Could we ‘obviously assume’ that before?
我们以前难道不能‘显然地假设’这一点吗?
Raymond Arnold: Up until now, I’ve been assuming the latest models are basically safe to use. I think we can no longer obviously assume that.
Raymond Arnold:直到最近,我一直假设最新的模型基本上是可以安全使用的。我认为我们现在不能再‘显然地假设’这一点了。
We are talking price. It is obviously safe on a personal level to use Astra or Fable 5.1 for ordinary chat tasks, indeed safer than using previous models.
我们在谈论价格。在个人层面上,使用 Astra 或 Fable 5.1 进行普通聊天任务显然是安全的,事实上比使用之前的模型更安全。
The question is, at what point are you worried about using Astra for agentic tasks, or giving Astra access to sufficient credentials to do serious damage, and how does that compare to our trust in Sol or Fable and so on.
问题是,在什么情况下你会担心使用 Astra 执行代理任务,或者授予 Astra 足以造成严重损害的权限?这与我们对 Sol 或 Fable 等模型的信任相比如何?
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力