AI周报187:Opus 5.5发布、美国反ASI法案与军事误报事件
AI #187: Coming Into Play
本周资讯密度极高,从旗舰模型更新到立法动态再到军事误报案例,全面覆盖AI行业核心进展,适合快速掌握全貌。
Opus 5.5 was released on Tuesday. I covered the system card yesterday, and will cover its capabilities soon. By all reports it is an excellent model. There are lots of fun videos going around that Opus has generated, which I will include as part of that.
Opus 5.5 于周二发布。我昨天介绍了系统卡片,很快将介绍其能力。据各方报道,这是一款出色的模型。网上流传着许多由 Opus 生成的有趣视频,我将把它们纳入其中。
OpenAI released a new cheaper and improved Sol and Luna. No one is talking about them due to Opus 5.5, but these should be an important upgrade under the hood.
OpenAI 发布了更便宜且改进后的 Sol 和 Luna。由于 Opus 5.5 的发布,没人谈论它们,但它们在底层应是一次重要的升级。
Bernie Sanders and Greg Casar have formally introduced the Ban Artificial Superintelligence Act. That means we get to read (RTFB) it. As always, I reserve judgment on particular bills until I can read them in detail. MIRI did so, and endorses the bill as directly confronting the extinction threat. I hope to do an RTFB soon.
伯尼·桑德斯(Bernie Sanders)和格雷格·卡萨尔(Greg Casar)正式提出了《禁止人工超级智能法案》。这意味着我们可以阅读(RTFB)该法案全文。一如既往,在我能详细阅读之前,我对具体法案保留判断。MIRI 也这样做了,并支持该法案,认为它直接应对了灭绝威胁。我希望很快能做一次 RTFB。
I have spun two things off the weekly:
我从每周内容中拆分出了两项:
- Coverage of the quest for the right embedded evaluators and related questions and attacks, which will become its own post.
- Some issues related to cooperative alignment, which may get folded into the model welfare post.
- 关于寻找合适嵌入式评估器及相关问题与攻击的报道,这将自成一篇博文。
- 一些与合作对齐相关的问题,可能会被合并到模型福利博文中。
I also might, in addition to a potential RTFB on the Sanders bill, do full podcast coverage of Jensen Huang on Ezra Klein, if time and emotion permit.
此外,如果时间和情绪允许,除了可能对桑德斯法案进行 RTFB 外,我还打算对 Jensen Huang 在 Ezra Klein 播客上的内容进行完整播客覆盖。
Otherwise, I have caught up on the news. To the extent the news allows there will be reduced posting (I know, I know) for the next few weeks as I race to complete another high impact project.
否则,我已经追上了新闻。只要新闻允许,未来几周发帖量会减少(我知道,我知道),因为我正全力完成另一个高影响力项目。
And yes, we are going to keep calling AI AI, and ASI ASI, thank you very much.
是的,我们将继续称 AI 为 AI,称 ASI 为 ASI,非常感谢。
Table of Contents
目录
- On The Terms Superintelligence and ‘Super Intelligence’. TYFYATTM.
- Language Models Offer Mundane Utility. Find the fraud in plain sight.
- Language Models Don’t Offer Mundane Utility. Risk an international incident.
- Language Models Can Only Work With What You Give Them. Stale intel.
- Huh, Upgrades. GPT-6 Sol and Luna, Grok 4.7, MiMo-v2.6-Pro.
- On Your Marks. The AIs exceed human baseline at Metaculus predictions.
- Get My Agent On The Line. Do not let Muse become your muse.
- Deepfaketown and Botpocalypse Soon. Literary critics have no taste.
- Fun With Media Generation. Knowledge of the future is unevenly distributed.
- Copyright Confrontation. NYT vs. OpenAI is unsealed. Some ugly details.
- Cyber Lack of Security. Mistral is hacked, allegedly. Others definitely hacked.
- Hugging the Face. OpenAI rogue agents hacked Australia.
- Hacking Into OpenAI. Three white hat dudes with Opus 5 hacked into OpenAI.
- They Took Our Jobs. No one has internalized what is coming.
- Get Involved. Derek Thompson is hiring a part-time social media producer.
- Anthropic Has a Wet Lab and a Potential Gene Editing Technique. Oh good.
- Introducing. Computer-10, Agi.fyi, Astra for Law.
- In Other AI News. Many at labs and UK AISI are burning out.
- Show Me the Money. Harvey gets good enough to use, ruins its margins.
- Bubble, Bubble, Toil and Trouble. AOC on whether AI labs are too big to fail.
- Anthropic Approaches Recursive Self-Improvement. Are you nervous yet?
- Others Approach Recursive Self-Improvement. OpenAI and z.AI as well.
- Burden of Proof. OpenAI is now sitting on over 100 mathematical results.
- Quickly, There’s No Time. We are remarkably close to on track for AI 2027.
- Left Wing Americans Really Hate AI For Different Reasons. Weird ones.
- Chip City. Andy Masley gets a piece in The Atlantic.
- Pick Up the Phone. There will be a new (virtual) phone to pick up.
- The Week in Audio. Jensen on Ezra Klein, me on Pushkin, Altman, Soares.
- People Just Say Things.
- Venkatesh Rao Stops Writing. Ideas will now be expressed via AI. I am sad.
- A Call for Control of Frontier AI Models. Many countries are signing.
- Calls For Pacing The Frontier. NYT Editorial board, Tao, Fukuyama.
- A Matter of Antitrust. OpenAI and Anthropic got close to a testing deal.
- A Matter of Liability. The AI labs should be liable for harms. They agree.
- Quest for Sane Regulations. Trump is open to AI guardrails.
- Rhetorical Innovation. People go back to living in the shadow of extinction.
- Tap the Sign. It’s a limited number of signs.
- A Matter of Some Debate. Who is willing or eager to debate who?
- Astra Is Hard to Monitor. This goes beyond its Chain of Thought.
- Anticipating What a Smarter Intelligence Can Do Is Impossible. People still ask.
- Would You Look At All These Goalposts. Why do they keep moving?
- I, Robot. Your bespoke robot software is no match for my bitter lesson.
- People Are Worried About AI Killing Everyone. Worried about being worried.
- Other People Are Not As Worried About AI Killing Everyone. It’s true.
- Joe Rogan. Let AI take over, he says, so we can end all wars. Well, yes, I guess…
- The Lighter Network Graph. Concerning.
- The Lighter Side. You would never be convinced by such mysterious means.
- 关于‘超级智能’与‘超智能’术语。TYFYATTM。
- 语言模型提供平庸效用。发现明目张胆的欺诈。
- 语言模型不提供平庸效用。引发国际事件的风险。
- 语言模型只能处理你给它的东西。过时情报。
- 嗯,升级。GPT-6 Sol 和 Luna、Grok 4.7、MiMo-v2.6-Pro。
- 各就各位。AI 在 Metaculus 预测上超越人类基线。
- 让我的 Agent 上线。不要让 Muse 成为你的缪斯。
- 深度伪造镇与机器人末日将至。文学评论家毫无品味。
- 媒体生成乐趣。对未来知识的分布不均。
- 版权对峙。纽约时报诉 OpenAI 案解封。一些丑陋的细节。
- 网络安全缺失。Mistral 据称被黑客入侵。其他公司肯定已被入侵。
- 拥抱脸书。OpenAI 失控 Agent 入侵澳大利亚。
- 入侵 OpenAI。三位白帽黑客使用 Opus 5 入侵了 OpenAI。
- 他们抢走了我们的工作。没有人真正意识到即将到来的一切。
- 参与其中。德里克·汤普森正在招聘兼职社交媒体制作人。
- Anthropic 拥有湿实验室和潜在的基因编辑技术。哦,太好了。
- 介绍。Computer-10、Agi.fyi、Astra for Law。
- 其他 AI 新闻。许多实验室员工和英国人工智能安全研究所(UK AISI)人员正面临倦怠。
- 让我看看钱。Harvey 变得足够好用,却毁掉了利润率。
- 泡泡,泡泡,辛劳与麻烦。AOC 探讨 AI 实验室是否大到不能倒。
- Anthropic 迈向递归自我改进。你感到紧张了吗?
- 其他公司也在迈向递归自我改进。OpenAI 和 z.AI 也是如此。
- 举证责任。OpenAI 现在手握超过 100 项数学成果。
- 快一点,没时间了。我们离实现 AI 2027 目标近在咫尺。
- 美国左翼出于不同原因真正讨厌 AI。奇怪的原因。
- 芯片之城。安迪·马斯利在《大西洋月刊》上发表文章。
- 拿起电话。将有一款新的(虚拟)电话供人接听。
- 本周音频节目。詹森接受伊兹拉·克莱因采访,我在 Pushkin 播客上谈论阿尔特曼和索雷斯。
- 人们只是随口说说。
- 文卡特什·劳停止写作。思想今后将通过 AI 表达。我很伤心。
- 呼吁控制前沿 AI 模型。许多国家正在签署协议。
- 呼吁为前沿技术发展设定节奏。《纽约时报》编辑委员会、陶哲轩、福山。
- 反垄断问题。OpenAI 和 Anthropic 曾接近达成一项测试协议。
- 责任问题。AI 实验室应对造成的伤害负责。他们对此表示同意。
- 寻求合理的监管。特朗普对 AI 护栏持开放态度。
- 修辞创新。人们重新生活在灭绝的阴影之下。
- 点击签名。签名数量有限。
- 存在争议的问题。谁愿意或急于与谁辩论?
- Astra 很难监控。这超出了其思维链的范畴。
- 预测更智能的智能体能做什么是不可能的。人们仍在追问。
- 看看这些不断移动的球门柱。为什么它们一直在变?
- 我,机器人。你定制的机器人软件抵不过我的痛苦教训。
- 人们担心 AI 会杀死所有人。他们甚至担心自己过于担忧。
- 其他人并不像那样担心 AI 会杀死所有人。这是事实。
- 乔·罗根。他说让 AI 接管,这样我们就能结束所有战争。嗯,是的,我想……
- Lighter 的网络图。令人担忧。
- Lighter 的另一面。你绝不会通过这种神秘的手段被说服。
On The Terms Superintelligence and ‘Super Intelligence’
关于术语“超级智能”和“Super Intelligence”
Donald Trump, not content with the Gulf of America and Lake America, has decided to proclaim that henceforth Artificial Intelligence will be called ‘Super Intelligence.’
唐纳德·特朗普,在拥有“美利坚海湾”和“美利坚湖”之后仍不满足,已决定宣布从此人工智能将被称为“Super Intelligence”。
No. This blog will continue to refer to AI as AI, or Artificial Intelligence.
不。本博客将继续把 AI 称为 AI,或 Artificial Intelligence。
Being the President does not actually give you the right to change how English works.
当总统实际上并不能赋予你改变英语运作方式的权力。
This is also the worst possible nightmare of a namespace clash he could have created.
这也是他可能制造的最糟糕的命名空间冲突噩梦。
If Trump had gone by his original poll choices, and decided to call it ‘Extreme Intelligence,’ ‘Superior Intelligence’ or ‘Supreme Intelligence,’ or perhaps ‘American Intelligence’ or ‘Trump Intelligence’ or whatever, it would be very funny tilting at windmills. I would make a lot of jokes about it.
如果特朗普遵循他最初的民意调查选择,并决定称其为“Extreme Intelligence”、“Superior Intelligence”或“Supreme Intelligence”,或者也许是“American Intelligence”或“Trump Intelligence”之类的,那将会非常有趣地对着风车冲锋。我会对此开很多玩笑。
Instead, Trump has chosen the term for ‘an AI better than everyone at everything.’
相反,特朗普选择了“在所有方面都比所有人更优秀的 AI”这一术语。
So now, users of both the term superintelligence and the term Super Intelligence will be saying that the AI in question is smarter than they are, but for different reasons.
因此,现在使用 superintelligence 和 Super Intelligence 这两个术语的用户都会说相关的人工智能比他们更聪明,但原因不同。
Thus, I am declaring that henceforth:
因此,我宣布从此以后:
If someone uses the term Super Intelligence (or ‘SUPER INTELLIGENCE’) as multiple words, or uses the initials SI, then this means AIMAGA, or ‘Artificial Intelligence that Makes America Great Again.’ SIMAGA is also acceptable.
如果有人使用 Super Intelligence(或‘SUPER INTELLIGENCE’)作为多个单词,或使用首字母缩写 SI,则这意味着 AIMAGA,即“让美国再次伟大的 AI”。SIMAGA 也是可以接受的。
If someone uses the term superintelligence (or Superintelligence) as a single word, or the phrase Artificial Superintelligence, or the initials ASI, then this refers to ‘Artificial Intelligence that does approximately all the things better than humans do.’
如果有人使用 superintelligence(或 Superintelligence)作为一个单词,或短语 Artificial Superintelligence,或首字母缩写 ASI,则这指的是“大致在所有方面都比人类做得更好的 AI”。
If you are speaking verbally, such that the pause might be imagined or missed, then for now either include the word ‘artificial’ or use the initials ASI, to avoid confusion.
如果你是在口头说话,以至于停顿可能被想象或错过,那么目前请包含单词“artificial”或使用首字母缩写 ASI,以避免混淆。
If you use the term AGI, or Artificial General Intelligence, no one will know what you mean, but that has nothing to do with Trump’s announcement. Carry on.
如果你使用 AGI,或 Artificial General Intelligence 这个术语,没人会知道你的意思,但这与特朗普的公告无关。继续吧。
Thank you for your attention to this matter.
感谢您对此事的关注。
Language Models Offer Mundane Utility
语言模型提供日常效用
Give anyone the ability to generate extensive reports on things like PPP fraud. We have not yet seen the impact that such information could have.
赋予任何人生成关于PPP欺诈等广泛报告的能力。我们尚未看到此类信息可能产生的影响。
Boost productivity in new materials R&D by an estimated 10x. The estimate is from Xiaomi, so treat it with skepticism, but I don’t doubt major acceleration. They attribute this to their MiMo-v2.6-Pro model, whereas I would say ‘even MiMo can give’ this kind of boost.
据估计,将新材料研发的 productivity 提升约10倍。该估算来自小米,因此请持怀疑态度看待,但我并不怀疑会有重大加速。他们将其归因于其 MiMo-v2.6-Pro 模型,而我认为‘即使是 MiMo 也能带来’这种程度的提升。
Language Models Don’t Offer Mundane Utility
语言模型不提供日常效用
Risk a major international incident by incorrectly telling the US military that a Chinese ship is transporting nuclear weapon components. Luckily, before the operation happened, they checked and realized the AI had misidentified the material. CNN calls this a ‘hallucination’ but this seems to me more like an error.
因错误告知美军一艘中国船只正在运输核武器组件,从而引发重大国际事件的风险。幸运的是,在行动发生前,他们进行了核查并意识到 AI 误判了材料。CNN 称此为‘幻觉’,但在我看来这更像是一个错误。
CNN also says this ‘almost started a war’ which is vastly overstating things. America and China are not about to go to full war over one mistakenly boarded ship.
CNN 还称此事‘几乎引发了战争’,这严重夸大了事实。美国和China 不会仅因一艘被误登的船只就全面开战。
Katie Bo Lillis, Zachary Cohen (CNN): “The internal tools are mostly just copies of the commercial stuff wearing lipstick,” a former senior US official familiar with the AI systems used by military and intelligence analysts [said].
Katie Bo Lillis, Zachary Cohen (CNN):‘内部工具基本上只是涂了口红(稍作修饰)的商业产品副本,’一位熟悉军事和情报分析师所用 AI 系统的前美国高级官员[表示]。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力