Opus 5 用起来为何感觉更差?
Opus 5 为何用起来感觉更差?
Why does Opus 5 feel worse to work with?
为什么 Opus 5 用起来感觉更差?
Posted on 2026-08-14 :: Source Code
发布于 2026-08-14 :: 源代码
Why does Opus 5 feel worse to work with?
为什么 Opus 5 用起来感觉更差?
In my opinion and that of the colleagues I've spoken with, working with Opus 5 feels like a downgrade compared to Opus 4.7, Opus 4.8, and Fable.
根据我和同事们的看法,使用 Opus 5 的感觉相比 Opus 4.7、Opus 4.8 和 Fable 是一种降级。
I'm not claiming a step backwards in capabilities – it is a more capable model than Opus 4.7 and Opus 4.8 and even rivals Fable in benchmarks, yet these other models feel better to work with. I believe this is because they:
我并不是说它在能力上退步了——它比 Opus 4.7 和 Opus 4.8 更强大,甚至在基准测试中能与 Fable 匹敌,但其他模型用起来感觉更好。我认为这是因为它们:
- stop and ask questions if my intent was unclear,
- don't make assumptions without checking,
- and don't reinterpret or update my plans without asking.
- 如果我的意图不明确,会停下来提问;
- 不会未经核实就做出假设;
- 并且不会未经询问就重新解释或更新我的计划。
Because of this, they don't require the careful babysitting that Opus 5 does.
正因为如此,它们不需要像 Opus 5 那样需要仔细的监督。
Baseless speculation
无根据的猜测
I suspect this is the result of two compounding forces at Anthropic, and in current frontier labs in general.
我怀疑这是 Anthropic 以及当前前沿实验室中两股力量共同作用的结果。
First, the desire to create a self-improving AI that is capable of recursively bootstrapping itself to AGI/ASI.
首先,是创造一种能够递归地自我提升到 AGI/ASI 的自我改进型 AI 的愿望。
Second, the pressure to score highly on benchmarks. Although it's an open secret that many benchmark tasks are ill-defined, unfair, hackable, or otherwise broken, a good benchmark task is self-contained. It can be solved. It doesn't require hints, reading the task creator's mind, or outside information to pass.
其次,是在基准测试中取得高分的压力。尽管许多基准测试任务定义不清、不公平、可被利用或存在其他问题,这是公开的秘密,但一个好的基准测试任务是自包含的。它可以被解决。它不需要提示、猜测任务创建者的想法或外部信息就能通过。
That doesn't mean a good task can only have one correct answer, just that it should score all unambiguously correct answers equally.
这并不意味着一个好的任务只能有一个正确答案,只是它应该对所有明确正确的答案给予同等的评分。
Selecting for models that do well on benchmarks (and indeed training for them or on RLVR tasks in general) inherently selects for models that make bold, usually-correct assumptions in the face of ambiguity. It penalizes models with a tendency to stop and ask for clarification or direction.
选择在基准测试中表现良好的模型(实际上也是针对它们进行训练,或普遍针对 RLVR 任务进行训练)本质上会选择那些在面对模糊性时做出大胆且通常正确的假设的模型。这会惩罚那些倾向于停下来询问澄清或方向的模型。
Unfortunately, that's exactly what most of us want from a coding agent.
不幸的是,这正是我们大多数人对编码代理的期望。
Try as you might, it's nearly impossible to get the entirety of the context, intentions, business implications, budget constraints, and what-have-you written down and accessible to a coding agent. There will invariably be ambiguity and choices to be made, and it is nice to know that an agent will stop and ask when needed.
尽管你尽力而为,但几乎不可能将全部上下文、意图、业务影响、预算限制以及诸如此类的信息都写下来并让编码代理获取。总会有模糊性和需要做出的选择,而知道代理会在需要时停下来询问是很好的。
Real life just isn't a benchmark. There isn't a guaranteed right answer to every question, nor even a set of right answers, and with real-life consequences on the line, I do not want an agent taking its best guess!
现实生活不是基准测试。每个问题并不都有保证的正确答案,甚至没有一组正确答案,而且由于涉及现实后果,我不希望代理凭最佳猜测行事!
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力