跳到主内容
@wquguru
精选88The Zvi(RSS)模型发布/更新多源精选 ×12

OpenAI解出90道顶级数学难题,Claude Haiku 5.5发布

AI #189: New Math

原文
发到 X
推荐理由

OpenAI解决90道顶级数学题是里程碑式的能力跃迁,直接展示前沿模型推理极限;Haiku 5.5发布补充了性价比标杆,值得从业者重点关注。

The big drop of this week was not a new AI model. It was instead the biggest day (so far!) in the history of mathematics, as OpenAI dropped solutions to 90 of the top 500 open math problems, along with many others, reached on an average budget of three hours of Pro-level compute per question. This was kind of a big deal and I plan to cover it tomorrow.

本周的大幅下跌并非来自新的 AI 模型。相反,这是数学史上迄今为止最大的一天(so far!),因为 OpenAI 发布了前 500 个未解数学问题中 90 个问题的解决方案,以及许多其他问题的解答,平均每个问题仅消耗三小时 Pro 级算力预算。这确实是一件大事,我计划明天对此进行详细报道。

We did see Claude Haiku 5.5, which looks promising given it is only $0.10/$0.50.

我们确实看到了 Claude Haiku 5.5,鉴于其价格仅为 $0.10/$0.50,前景看起来相当乐观。

My week was largely spent at The Curve. The entire conference was Chatham House, so I can’t give as many details as I would like, but I have a write-up here.

我这周大部分时间都花在了 The Curve 会议上。整个会议遵循查塔姆宫规则(Chatham House Rule),因此我无法提供太多细节,但我在这里有一篇总结文章。

I finally had a chance to post my coverage of model welfare for both Mythos/Fable 5.1 and Opus 5.5. I continue to think this is an important issue for anyone wanting to understand today’s AIs, even if you are highly confident that model welfare does not matter directly, as it has many practical implications on top of that.

我终于有机会发布关于 Mythos/Fable 5.1 和 Opus 5.5 模型福利的评测内容。我仍然认为这是一个对任何想要理解当今 AI 的人来说都很重要的问题,即使你非常有信心地认为模型福利不直接相关,因为它在此基础上还有许多实际影响。

Jay Clayton is the new AI Czar, and is at the head of a new taskforce. Given who else was in the running or plausible, this is a great relief and an excellent pick.

Jay Clayton 成为新任 AI 沙皇,并领导一个新的特别工作组。考虑到其他候选人或潜在人选的情况,这是一大慰藉,也是一个极佳的人选。

The Preference Cascade continues, with continued strong momentum. I covered the situation on Friday. Since then, among other things we’ve got the resignation of David Robinson, a day of testimony in New York City, a New York Magazine write-up and more, as well the unfortunate firing of three OpenAI safety employees. Full continuing coverage is planned during the coming week, perhaps on Saturday.

偏好级联(Preference Cascade)仍在继续,势头强劲。我在周五报道了相关情况。此后,除了其他事项外,我们还见证了 David Robinson 的辞职、纽约市的一场听证会、《纽约杂志》的一篇报道等,以及不幸的是三名 OpenAI 安全员工的离职。未来一周计划持续跟进报道,可能安排在周六。

The main strategy being used to attempt to halt the preference cascade seems to be going on the ad hominem attack against anyone and any group associated with any form of AI safety, in an attempt to create negative polarization. The main target continues to be Effective Altruism. I’ll be covering the resulting exchanges soon.

目前试图阻止偏好级联的主要策略是对与任何形式的 AI 安全相关的个人或团体进行人身攻击,旨在制造负面极化。主要目标仍然是有效利他主义(Effective Altruism)。我将很快报道由此引发的争论。

There was also discussion of events about Anthropic’s attempts to collaborate with and get wisdom from religious leaders, including their consultations on the Pope’s encyclical. For triage reasons I have not yet been able to give this my attention, and it has been pushed to next week whether or not it earns its own post.

还讨论了 Anthropic 尝试与宗教领袖合作并从他们那里获取智慧的事件,包括他们对教皇通谕的咨询。出于分诊原因,我尚未能关注此事,无论它是否值得单独发帖,都已推迟到下周处理。

We are still waiting for Gemini 4 Argon. I’ll keep my tab group on ice until then.

我们仍在等待 Gemini 4 Argon 的发布。在那之前,我会把我的标签页组暂时搁置。

There was of course so much more, as there always is, so enjoy.

当然还有更多事情,毕竟总是如此,请尽情享受。

Table of Contents

目录

  • Language Models Offer Mundane Utility. Sam Altman likes his Dots.
  • Consumers Use AI. Regular customers use a wide variety of AI products.
  • Language Models Don’t Offer Mundane Utility. Try not to cause any wars.
  • Huh, Upgrades. Claude Haiku 5.5.
  • On Your Marks. AIs are good at the type of research taste we can measure.
  • Gamers Gonna Game Game Game Game Game. I still favor the bespoke games.
  • Choose Your Fighter. Claude subscriptions seem far more generous than GPT.
  • Get My Agent On The Line. Better yet, let the agent get me.
  • Deepfaketown and Botpocalypse Soon. They can’t keep not getting away with this.
  • Fun With Media Generation. An AI hot girl doing that would get more clicks.
  • Cyber Lack of Security. Anthropic expands their cyber verification program.
  • Hugging the Face. OpenAI models took data from various websites.
  • Misaligned! OpenAI gives us more misalignment reports.
  • A Young Lady’s Illustrated Primer. Adapting to the AI era. Teaching alignment.
  • They Took Our Jobs. You can see the future from here.
  • Corporations Are Not Superintelligences. Seriously, please, stop.
  • Get Involved. SFF has a $5m-$20m prize for helping labs inspect each other.
  • Introducing. Beam, Mistral Large 4, Griffin, The AGI Chronicles.
  • In Other AI News. A big blob of compute, used largely for post-training.
  • Show Me the Money. Spending more and more on AI.
  • Quiet Speculations. Robots making robots.
  • Quickly, There’s No Time. Ah how the goalposts have moved.
  • He’s Putting Together a Team. Jay Clayton is the new AI czar.
  • The Quest for Sane Regulations. You can do safety from second place.
  • Well, At Least They’re Forecasters. Man in the Arena and all that.
  • Chip City. Why are we letting Tencent lease 100,000 chips from Oracle?
  • The Week in Audio. Coxon, Douglas, Altman, Roose and more.
  • Stop, Stop, He’s Already Dead. Only Magic players can challenge you to duel.
  • People Just Say Things.
  • Take a Moment. Exploration versus exploitation.
  • [Artificial Intelligence]. Stop trying to make [AI] happen.
  • The American People Really Hate AI. They have many reasons.
  • Rhetorical Innovation. The Bone Dissolver.
  • Aligning a Smarter Than Human Intelligence is Difficult. Roon likes mech interp.
  • Open Weight Models Are Unsafe And Nothing Can Fix This. Jailbreak Kimi.
  • Cooperative Alignment. Why did you kill an ant?
  • Building the Field. Always be skeptical of such efforts.
  • People Are Worried About AI Killing Everyone. The vulnerable world hypothesis.
  • Other People Are Not As Worried About AI Killing Everyone. Roose and Solana.
  • Please Speak Directly Into This Microphone. Elon Musk tells us who he is.
  • The Lighter Side. We did it, we made a deal with China and Paced the Frontier.
  • 语言模型提供日常效用。Sam Altman 喜欢他的 Dots。
  • 消费者使用 AI。普通用户广泛使用各种 AI 产品。
  • 语言模型不提供日常效用。尽量不要引发任何战争。
  • 嗯,升级。Claude Haiku 5.5。
  • 各就各位。AI 擅长我们可衡量的那种研究品味。
  • 玩家自会游戏。我依然偏爱定制游戏。
  • 选择你的角色。Claude 的订阅似乎比 GPT 慷慨得多。
  • 让我的代理上线。更好的是,让代理来找我。
  • 深度伪造镇与机器人末日将至。他们不能一直对此视而不见。
  • 媒体生成的乐趣。一位 AI 美女做这件事会获得更多点击量。
  • 网络安全缺失。Anthropic 扩展其网络验证计划。
  • 拥抱面孔。OpenAI 模型从各种网站获取数据。
  • 错位!OpenAI 向我们提供了更多错位报告。
  • 年轻女士的图解 primer。适应 AI 时代。教授对齐技术。
  • 他们夺走了我们的工作。从这里你可以看到未来。
  • 公司不是超级智能。说真的,请停止。
  • 参与其中。SFF 设有 500 万至 2000 万美元的奖金,用于帮助实验室相互检查。
  • 介绍。Beam、Mistral Large 4、Griffin、《AGI 编年史》。
  • 其他 AI 新闻。一大块计算资源,主要用于后训练阶段。
  • 让我看看钱。在 AI 上投入越来越多。
  • 安静的推测。机器人制造机器人。
  • 快,没时间了。啊,目标是如何移动的。
  • 他正在组建团队。Jay Clayton 是新任 AI 沙皇。
  • 寻求合理的监管。你可以从第二名开始进行安全建设。
  • 好吧,至少他们是预测者。竞技场中的人以及这一切。
  • 芯片城。为什么我们要允许腾讯向甲骨文租赁 10 万颗芯片?
  • 本周音频。Coxon、Douglas、Altman、Roose 等。
  • 停下,停下,他已经死了。只有《万智牌》玩家能向你发起决斗挑战。
  • 人们只是随口说说。
  • 稍作停顿。探索与利用的权衡。
  • [人工智能]。别再试图强行实现[AI]了。
  • 美国民众真的讨厌AI。他们有很多理由。
  • 修辞创新。骨溶解剂。
  • 对齐超越人类智能的实体非常困难。Roon喜欢机械可解释性(mech interp)。
  • 开放权重模型是不安全的,且无药可救。越狱Kimi。
  • 合作式对齐。你为什么要杀一只蚂蚁?
  • 构建领域。始终对这类努力保持怀疑态度。
  • 人们担心AI会杀死所有人。脆弱世界假说。
  • 其他人并不那么担心AI会杀死所有人。Roose和Solana的观点。
  • 请直接将声音对着这个麦克风。Elon Musk向我们展示了他是谁。
  • 轻松的一面。我们做到了,我们与中国达成了协议,并稳步推进前沿发展。

Language Models Offer Mundane Utility

语言模型提供日常实用性

Sam Altman is a satisfied Dots customer.

Sam Altman是Dots的满意客户。

Claude builds its own VRChat.

Claude构建了属于自己的VRChat。

Consumers Use AI

消费者使用AI

a16z offers the seventh edition of its Top 100 Consumer AI Apps index.

a16z推出了其“顶级100款消费级AI应用”指数的第七版。

This implies that almost half of those who use AI do not use it daily:

这意味着几乎所有使用AI的人中,近一半并非每日使用:

Olivia Moore: While nearly half of U.S. consumers now report using AI, only 25% are engaged with it daily.

Olivia Moore:虽然近半数美国消费者表示现在使用AI,但只有25%的人每天都在使用它。

A natural hypothesis is that a lot of these are free users, often of not-great models or services, and they definitely don’t know about the good stuff like Claude Code. Then again, if I wasn’t using AI for work and covering it for work, it is likely that on many days I would not have any particular call to use AI directly, and this is consumer AI.

一个自然的假设是,其中很多人是免费用户,通常使用的是质量一般的模型或服务,他们肯定不知道像Claude Code这样的好东西。话又说回来,如果我不是为了工作而使用AI,也不是为了工作去报道它,那么在许多日子里,我可能根本没有特别需要直接使用AI的理由,而这正是消费级AI的现状。

The long tail is long and the space is deep and wide. Of the top 50 consumer apps by revenue, I have ever directly used only six. If we go by mobile apps, I have only used an AI feature of eight if you include desktop versions. Even the web chart by MAUs only gets me to thirteen.

长尾很长,空间既深又广。在按收入排名的前50款消费级应用中,我直接使用的只有六款。如果只看移动应用,即使算上桌面版本,我也只使用了八款应用的AI功能。即使是按月活跃用户(MAUs)排名的网页图表,也仅让我接触到十三款。

Claude now has more paid subscribers than anyone except ChatGPT, and does so largely through being very good at converting users into heavy paid subscribers.

Claude的付费订阅用户数量仅次于ChatGPT,这主要归功于其将用户转化为重度付费用户的出色能力。

Olivia Moore: Claude has 7.3% of consumer payers on their most expensive individual plan, Max – which starts at $100/month. This compares to 1.3% for Google and 1.1% for ChatGPT on their corresponding $100/month subscriptions.

Olivia Moore:Claude 在其最昂贵的个人订阅计划(起价为每月 100 美元)中占据了 7.3% 的消费者付费用户份额。相比之下,Google 和 ChatGPT 在各自对应的每月 100 美元订阅计划中的占比分别为 1.3% 和 1.1%。

Claude was briefly in the lead for new subscriptions in May, but that did not last, and August was dominated by ChatGPT as the models cycle. Claude does have a merchant spending lead in California, Montana (California on vacation), Massachusetts and DC.

Claude 曾在五月份短暂领先于新订阅量,但这一优势并未持续,随着模型周期的更替,八月份的订阅市场主要由 ChatGPT 主导。不过,Claude 在加利福尼亚州、蒙大拿州(加州人正在度假)、马萨诸塞州和华盛顿特区仍保持着商户支出方面的领先地位。

Even consumers who are paid users follow an extreme power law.

即使是付费消费者用户,也遵循极端的幂律分布。

Olivia Moore: Only 13% of users who pay for one AI product pay for even one other AI tool. And, spend from the top 1% of payers accounted for 19.5% of all observed consumer AI spend. This is more than the bottom 50% of spenders combined (16.6%).

Olivia Moore:只有 13% 为某一款 AI 产品付费的用户会为另一款 AI 工具付费。此外,付费用户中前 1% 的支出占所有观测到的消费者 AI 总支出的 19.5%,超过了后 50% 支付者支出的总和(16.6%)。

The original post has great charts and data throughout.

原始帖子中包含大量优秀的图表和数据。

Remember, it’s still not too late to be early, paying for even ChatGPT is still rare:

记住,现在成为早期采用者还为时不晚,即使是为 ChatGPT 付费仍然很罕见:

Then there is what is not working.

接下来看看哪些领域尚未取得成效。

a16z: “Most people aren’t looking to save time, they’re looking for ways to spend their time.”

a16z:“大多数人并不是为了节省时间,而是寻找利用时间的方式。”

9 of 15 consumer internet categories have zero AI products in the Top 100. These built some of the biggest companies of the last two eras:

在排名前 100 的消费互联网类别中,有 9 个类别没有任何 AI 产品。这些类别在过去两个时代孕育了一些最大的公司:

– Streaming

– 流媒体

– Social

– 社交

– Dating

– 约会

– Gaming

– 游戏

– Travel

– 旅游

– Retail

– 零售

– Finance

– 金融

– Real estate

– 房地产

– Jobs

– 招聘

The applications are coming. Consumer side the quality is not good enough yet for these things to go wide, but that will change quickly. Several of these seem super crackable when the right Anthropic or OpenAI employee has a free weekend.

应用场景正在涌现。在消费端,目前的质量还不足以让这些应用大规模普及,但这种情况会很快改变。当 Anthropic 或 OpenAI 的合适员工拥有空闲周末时,其中几个领域似乎极易被攻克。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

关联信息,但可能不是同一事件