Anthropic实验:AI多智能体协作反不如单模型,Jevons悖论未显
🔮 Why one AI is better than four #598
Agent协作中的群体思维陷阱是核心工程难题,这份一手实验数据直接挑战了多模型集成的直觉,值得所有做Agent架构的同学参考。
Good morning!
早上好!
We are looking for an outstanding economist to join us as an AI Economy Research Fellow. If you know someone we should speak to, send them our way.
我们正在寻找一位杰出的经济学家,加入我们要担任 AI 经济研究研究员。如果你知道我们应该联系的人,请把他们推荐给我们。
Great minds think (a little too much) alike
英雄所见略同(想得有点太多了)
A few months ago, we (alongside ) looked at whether AI is immune to groupthink. The answer was no. Blending several models’ answers kept about a quarter of the good ideas that had come from a single model. This is called the hidden-profile problem: when groups discuss what everyone already knows and don’t get to the knowledge that only one member holds. Anthropic has now run that classic experiment on agents: four agents must arrive at a decision. The evidence they hold in common points to the wrong option, while only a few agents (or just one) have the facts that lead to a correct decision. Getting it right means a small set of agents pressing its private facts and the others trusting them over the apparent consensus. After discussion, most model families chose correctly in only 17-36% of runs, while a single agent handed the entire evidence base got it right nearly every time. Only one model (somewhat) escaped: Mythos 5, at about 85% (why, we don’t know).
几个月前,我们(与 )一起探讨了 AI 是否对群体思维免疫。答案是否定的。融合多个模型的答案保留了来自单个模型的大约四分之一的优秀想法。这被称为隐藏档案问题:当群体讨论每个人都已经知道的事情,而没有接触到只有个别成员掌握的知识时。Anthropic 现在在智能体上运行了那个经典实验:四个智能体必须做出决策。他们共有的证据指向错误的选项,而只有少数智能体(或仅一个)拥有导致正确决策的事实。做对意味着一小部分智能体坚持其私有事实,而其他智能体信任它们而不是表面的共识。经过讨论,大多数模型家族仅在 17-36% 的运行中做出了正确的选择,而单个智能体掌握了全部证据几乎每次都做对了。只有一个模型(某种程度上)逃脱了:Mythos 5,达到了大约 85%(为什么,我们不知道)。
I see two problems at work here. First, LLMs lack diversity (they are low-variance): set 30 agents the same coding task and 18 of them will name their git branch identically. Second, agents lack the institutions that make human groups robust: reputation, recourse and protection for the lone dissenter. These aren’t necessarily unfixable, but it’s not yet clear what the fix is. On the diversity side, I particularly like the solutions Thinking Machines puts forward: an ecosystem of AIs raised in different places, with different values and purposes, “keeping the weirdness alive.” After all, most good ideas started weird.
我在这里看到了两个问题。首先,LLM 缺乏多样性(它们是低方差的):给 30 个智能体相同的编码任务,其中 18 个会以相同的方式命名他们的 git 分支。其次,智能体缺乏使人类群体具有韧性的制度:声誉、救济措施以及对唯一异议者的保护。这些不一定无法解决,但目前还不清楚解决方案是什么。在多样性方面,我特别喜欢 Thinking Machines 提出的解决方案:一个在不同地方成长、拥有不同价值观和目的的 AI 生态系统,“保持怪异感”。毕竟,大多数好主意起初都是怪异的。
When will the Jevons paradox kick in?
杰文斯悖论何时会显现?
In our State of AI report, we found a positive but underwhelming elasticity for tokens. A 10% price cut lifts token use by 12–18%: enough to raise total spend, but not by much.
在我们的《AI 现状报告》中,我们发现 token 的弹性为正但令人失望。价格降低 10%,token 使用量增加 12–18%:足以提高总支出,但幅度不大。
Patrick Saner made a comment that made me rethink why: “the cost per token is irrelevant. What matters is the cost of completing a useful unit of work.” Elasticity might be underwhelming because users haven’t found a way to properly price “a useful unit of work.” Firms exist exactly to avoid pricing work. Especially for knowledge work, we buy a lot of it in bundles: a salary, a retainer, an hour. Creating a priceable task from knowledge work is not easy. Some may have found a useful unit: since October 2023 the top 1% of firms raised AI spend per employee by $6,542. The median rose only $9.63. I would guess this is mostly software, where AI is both most proven and, in a sense, most measurable (commits, pull requests and releases).
Patrick Saner 发表了一条评论,让我重新思考了原因:“每个 token 的成本无关紧要。重要的是完成一个有用工作单位的成本。”弹性可能显得不足,因为用户尚未找到为“一个有用的工作单位”正确定价的方法。企业的存在正是为了避免对工作定价。尤其是对于知识型工作,我们通常以捆绑形式购买大量此类服务:如薪资、固定费用或按小时计费。将知识型工作转化为可定价的任务并不容易。有些人可能找到了一个有用的单位:自 2023 年 10 月以来,排名前 1% 的企业使每位员工的 AI 支出增加了 6,542 美元。中位数仅上升了 9.63 美元。我猜测这主要涉及软件领域,在那里 AI 既最具验证性,又在某种意义上最可衡量(如代码提交、拉取请求和发布)。
Read more
阅读全文
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力