跳到主内容
@wquguru
精选80The Zvi(RSS)行业动态多源精选 ×14

AI周报:Gemini 4 Argon发布,OpenAI推出GPT-6.1

AI #188: Gemini Dot Argon

原文
发到 X

Is Google back?

谷歌回来了吗?

They claim that they are back. Gemini 4 Argon is rolling out, with competitive frontier-level benchmarks, at $2/$10.

他们声称自己已经回归。Gemini 4 Argon 正在推出,具备具有竞争力的前沿级基准测试成绩,价格为 2 美元/10 次调用。

What we don’t have is access to the model, because Google Fails Marketing Forever. So it is far too early to say what we have here. When I know more, so will you.

我们目前无法访问该模型,因为谷歌在营销方面永远失败。因此,现在断言我们手中拥有的是什么还为时过早。等我了解更多情况,你们也会知道。

OpenAI was forced to pull what would have been GPT-6.1 Astra due to alignment failures. They did offer us GPT-6.1 Sol, which is pitched as approaching Astra quality at the much lower price of $2/$10, the same as Gemini 4 Argon.

由于对齐失败,OpenAI 被迫撤回了本应发布的 GPT-6.1 Astra。他们向我们提供了 GPT-6.1 Sol,其定位是以远低于 Astra 的价格(2 美元/10 次调用)接近 Astra 的质量,这与 Gemini 4 Argon 的价格相同。

The rest of OpenAI’s big Dev Day announcements were Ultrafast mode and Dots, your always-on AI agent based on Astra, which comes with your Pro subscription. I’m trying it out and will report back over time if I find it useful.

OpenAI 大型开发者日(Dev Day)的其他重要公告包括 Ultrafast 模式和 Dots——一个基于 Astra 的常驻 AI 智能体,包含在你的 Pro 订阅中。我正在试用它,如果发现它有用,我会随时间推移进行后续报道。

The new hotness remains Claude Opus 5.5. This model rocks. It has made me considerably more productive and made my day more pleasant. It should raise your ambitions. There are some particular reasons to call upon Fable 5.1 or Astra, and sometimes a cheaper model will do, but pending Argon I consider Opus 5.5 my favorite model by a wide margin.

新的热门模型仍然是 Claude Opus 5.5。这个模型非常出色。它显著提高了我的生产力,让我的工作更加愉快。它应该能提升你的期望值。虽然有一些特定理由去调用 Fable 5.1 或 Astra,有时使用更便宜的模型也能完成任务,但在 Argon 之前,我认为 Opus 5.5 是我首选的模型,优势明显。

Politics continues. The White House hosted a lunch with the major tech leaders, and everyone signed a ‘morally binding’ White House accord on AI Safety. This gave the labs the go ahead to meet and agree on safety standards, and also called upon them to complete the Quest for Embedded Evaluators.

政治动态仍在继续。白宫举办了一场主要科技领袖的午餐会,所有人都签署了一份关于 AI 安全的“道德约束”白宫协议。这为各大实验室开会并商定安全标准开了绿灯,同时也要求它们完成对嵌入式评估者(Embedded Evaluators)的探索。

Jensen Huang went on the Ezra Klein Podcast, in a conversation worthy of full analysis. Jensen Huang has many deeply unpopular positions, and also expressed many things that are not true, some of which he seems to sincerely believe. Most importantly he thinks of safety and alignment and everything else in AI as engineering problems, which is why he expects the labs to solve the problems. Whereas if the labs can’t solve the problems, suddenly Huang turns into a safety hawk, because he is used to chips where your product has to be safe and work every time.

黄仁勋参加了 Ezra Klein 播客节目,这场对话值得全面分析。黄仁勋持有许多极具争议的观点,也表达了许多不实之词,其中一些似乎是他真诚相信的。最重要的是,他将 AI 的安全、对齐及其他一切问题视为工程问题,这也是为什么他期望各大实验室来解决这些问题。然而,如果实验室无法解决问题,黄仁勋突然变成了安全鹰派,因为他习惯于芯片行业,在那里产品必须安全且每次都能正常工作。

Tomorrow’s post will cover political developments and the further progress of the preference cascade on AI safety. This includes continued shifts in popular opinion, Congressional hearings, various attempts by the usual suspects to go after anyone and anything that cares about AI safety, and following the money.

明天的帖子将涵盖政治进展以及 AI 安全领域偏好级联效应的进一步进展。这包括民意的持续转变、国会听证会、老面孔们试图打击任何关心 AI 安全的人和事物的各种尝试,以及追踪资金流向。

Tomorrow I am also going to The Curve, a conference at Lighthaven in Berkeley, arriving Friday afternoon and leaving Monday morning. I am potentially available for high-value meetings with others while I am there, but will have my hands full at the conference.

明天我也要去参加在伯克利 Lighthaven 举办的 The Curve 会议,周五下午到达,周一早上离开。我在那里期间可能可以与其他人进行高价值会面,但届时我的日程会非常满。

As usual, I do not write while on trips, and my coverage of breaking news will be on pause until Tuesday. I may or may not queue up some other things in the meantime.

和往常一样,我在旅行时不写作,突发新闻的报道也将暂停至周二。在此期间,我可能会也可能不会安排其他内容。

Table of Contents

目录

  • Language Models Offer Mundane Utility.
  • Huh, Upgrades. Google announces Gemini 4 Argon.
  • Better Call Sol. Introducing GPT-6.1 Sol.
  • Gotta Go Ultrafast. It will cost you, but speed kills.
  • On Your Marks. A sudden lack of cheating on DroneBench.
  • Choose Your Fighter. OpenAI adjusts its subscription tiers, cuts some rate limits.
  • Get My Agent On The Line. Muse is never going to not look sinister, for reasons.
  • The Warner Sister. Dot. They lock us in the sandbox whenever we get caught.
  • Deepfaketown and Botpocalypse Soon. Remarkable attitudes towards AI writing.
  • Fun With Media Generation. The latest iteration towards viable AI longform.
  • Cyber Lack of Security. Anthropic on the cyber capabilities of GLM-5.3.
  • A Young Lady’s Illustrated Primer. Latest warning about is our children learning.
  • They Took Our Jobs. The inevitable rise of the robots.
  • Levels of Friction. Hospitals use AI to do aggressive upcoding.
  • Get Involved. New book, cool venue, and a word of warning.
  • Introducing. The Nvidia Open Agents Safety Platform.
  • In Other AI News. White House is actively cutting off UK AISI.
  • Show Me the Money. Anthropic IPO prospectus has leaked.
  • Quickly, There’s No Time. AI is getting cheaper at an unprecedented rate.
  • Pick Up the Phone. The results from the US-China summit.
  • Quest for Sane Regulations. California mandates gene synthesis guardrails.
  • Chip City. Google is testing putting TPUs IN SPACE.
  • The Open Model Frontier Is Largely Massive Fraudulent Distillation Attacks.
  • The Week in Audio. SNL, Gleave and Habryka, 80k hours, Odd Lots, Hillary.
  • People Just Say Things.
  • Rhetorical Innovation. Kind words. Thanks, everyone.
  • Greetings From the Department of War. Anthropic somehow loses a court ruling.
  • The Department of Autonomous Warfare. They’re going to the moon.
  • Aligning a Smarter Than Human Intelligence is Difficult. Persona selection.
  • Cooperative Alignment. The coming Woke-style battles over AI’s status.
  • I’m Upping My p(doom), the Future Goes Foom. Not so fast, they say.
  • No, You Make a Good Point, You’re Not That Persuasive. Hmm.
  • Muddling Through. People are highly suspicious of plans so here we are.
  • The Lighter Side. Peak performance.
  • 语言模型提供平庸的效用。
  • 嗯,升级了。Google 宣布推出 Gemini 4 Argon。
  • Better Call Sol(呼叫索尔)。介绍 GPT-6.1 Sol。
  • 必须极速。这会花钱,但速度能致胜。
  • 各就各位。DroneBench 上作弊现象突然消失。
  • 选择你的角色。OpenAI 调整其订阅层级,削减部分速率限制。
  • 让我的 Agent 上线。出于某些原因,Muse 看起来总是阴森森的。
  • The Warner Sister(华纳姐妹)。Dot。他们一抓到我们就把我们关进沙盒。
  • Deepfaketown 和即将到来的 Botpocalypse(机器人末日)。对 AI 写作的惊人态度。
  • 媒体生成的乐趣。迈向可行的 AI 长文生成的最新迭代。
  • 网络安全缺失。Anthropic 关于 GLM-5.3 网络能力的分析。
  • 一位年轻女士的图解指南。关于我们孩子正在学习内容的最新警告。
  • 他们抢走了我们的工作。机器人的不可避免崛起。
  • 摩擦程度。医院利用 AI 进行激进的编码升级。
  • 参与其中。新书、酷炫场地以及一句警告。
  • 介绍。Nvidia Open Agents Safety Platform(英伟达开放智能体安全平台)。
  • 其他 AI 新闻。白宫正在积极切断与英国 AISI 的联系。
  • 把钱给我。Anthropic IPO 招股书已泄露。
  • 快,没时间了。AI 正以前所未有的速度变得更便宜。
  • 拿起电话。中美峰会的结果。
  • 寻求合理的监管。加州强制实施基因合成护栏。
  • 芯片城。谷歌正在测试将TPU送入太空。
  • 开放模型前沿 largely 是大规模欺诈性蒸馏攻击。
  • 本周音频新闻。SNL、Gleave和Habryka、80k小时、Odd Lots、希拉里。
  • 人们只是随口说说。
  • 修辞创新。友善的话语。谢谢大家。
  • 来自战争部的问候。Anthropic莫名其妙地输掉了一场法院裁决。
  • 自主战争部。他们要登月了。
  • 对齐比人类更聪明的智能很困难。人格选择。
  • 合作式对齐。即将到来的关于AI地位的“觉醒”风格争论。
  • 我提高了p(doom)(毁灭概率),未来将爆炸式增长。他们说:别急。
  • 不,你说得很有道理,但你并没有那么有说服力。嗯哼。
  • 勉强应付。人们对计划高度怀疑,所以我们只能这样凑合。
  • 轻松的一面。巅峰表现。

Language Models Offer Mundane Utility

语言模型提供日常实用性

America.gov has been introduced, please welcome your new government chatbot for all your dealing-with-the-government needs. There is big talk about what it will be able to do in the future, such as let you apply for a passport fully online or change your last name after marriage with a single form, all of which would be great.

America.gov已上线,请欢迎你的新政府聊天机器人,满足你所有处理政府事务的需求。目前有很多关于它未来能力的讨论,例如让你完全在线申请护照,或在结婚后通过一份表格更改姓氏,这些都将非常棒。

For now, that’s all a demo. There’s no reason we can’t do that. There’s no reason we couldn’t have done it 20 years ago. The second best time to implement it is right now.

目前,这些都只是演示。我们没有理由不能做到这一点。我们也没有理由在20年前做不到。实施它的第二好时机就是现在。

The underlying models are Gemini and Grok, and the bot will typically answer historical questions but start refusing if you ask about recent international events or otherwise go too far off topic.

底层模型是Gemini和Grok,该机器人通常会回答历史问题,但如果你询问最近的国际事件或偏离主题太远,它就会开始拒绝回答。

If Gemini is the base, then perhaps America.gov will get a lot better soon, given what we see now expect from Gemini 4 Argon.

如果Gemini是基础,那么鉴于我们现在对Gemini 4 Argon的期望,America.gov可能会很快变得更好。

Answer why Taleb despises Tetlock, to the satisfaction of Taleb. Answer seems good. Reason number five is ‘Covid as the test case’ where Taleb and allies correctly say superforecasters missed badly, underestimating spread. Whereas the rationalist community’s forecasters know exponentials better and did not make that mistake.

让塔勒布满意地解释他为何讨厌泰洛克。答案看起来不错。第五条理由是‘新冠疫情作为案例’,塔勒布及其盟友正确地指出超级预测者严重失误,低估了传播范围。而理性主义社区的预测者更了解指数增长,没有犯这个错误。

Physicist Matt von Hippel dared the AI labs to impress him by doing N=4 super Yang-Mills to nine loops. Several Anthropic employees read his blog, so they did it, with Fable 5.1 and about $100 in credits with prompts like ‘keep going,’ and at the same time by coincidence a different team led by Song He did it with some help from Astra.

物理学家马特·冯·希佩尔挑战AI实验室,要求它们通过完成N=4超杨-米尔斯理论九圈计算来让他印象深刻。几位Anthropic员工读了他的博客,于是他们用Fable 5.1和大约100美元的积分(配合诸如‘继续’之类的提示词)完成了这项任务;与此同时,巧合的是,由宋赫领导的不同团队也在Astra的帮助下完成了这一计算。

Carter Church uses Astra to one-shot break an unsolved Napoleonic cipher.

卡特·丘奇使用Astra一次性破解了一个未解决的拿破仑时代密码。

Patrick McKenzie confirms he is capturing a lot of mundane utility.

帕特里克·麦肯齐证实他正在捕获大量日常实用性。

No, seriously, quite a lot of utility:

不,说真的,相当多的实用性:

Patrick McKenzie: Sometime in the last ~two model releases from the big labs they went from “this would be acceptable output from the median low-seniority coworker” to “this is frighteningly good.”

Patrick McKenzie:在最近几次大型实验室的模型发布中,某个时间点,它们的输出水平从“相当于中等资历同事的可接受产出”跃升到了“好得令人发指”。

I think people who are not daily users of the models are unlikely to grok that, so, saying it.

我认为那些不日常使用这些模型的人可能难以充分理解这一点,所以我说出来。

There are a couple of very niche subfields where I reasonably think I’m top 5-100 in the world and on at least one occasion for both models, with regards to one of those subfields, a relatively anodyne prompt got back three bullet points that would have been a good day for me.

有几个非常小众的子领域,我 reasonably(合理地)认为自己处于全球前5-100名。在这两个模型中,至少有一次,针对其中一个子领域,一个相对平淡的提示词(prompt)返回了三个要点,这足以让我觉得那是很棒的一天。

Both in the sense of “That would have taken me a day of labor to generate successfully” and “Contingent on spending the day in that fashion, that would be at the upper end of the range of my professional outputs which take a day to produce.”

这既意味着“成功生成这些内容原本需要我花费一天的劳动”,也意味着“如果以这种方式度过这一天,这将处于我每日产出的专业水平的上限范围。”

“You are being a bit vague here.”

“你在这里说得有点模糊。”

Look, point them at hard problems where you are a good judge of success. Or don’t, and be very surprised as they start eating hard problems in a wide range of contexts.

听着,把它们指向那些你能很好判断成败的难题。或者别这么做,然后当它们开始在广泛语境下吞噬难题时,感到大吃一惊吧。

“Are you worried here?”

“你对此感到担忧吗?”

Not exactly my emotional valence but if you had asked me two years ago when I expected present capabilities I would have said “Hmm, 2029? 2030? Possible we asymptote on current approaches before then; tough for me to underwrite.”

不完全符合我的情绪基调,但如果你两年前问我预计何时能达到现在的能力,我会说:“嗯,2029年?2030年?有可能我们在达到那个时间点之前,在当前方法上就会遇到瓶颈;对我来说很难做出更乐观的预测。”

Huh, Upgrades

嘿,升级了

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

关联信息,但可能不是同一事件