跳到主内容
@wquguru
精选85The Zvi(RSS)模型发布/更新

Claude Opus 5 能力接近 Fable,非 Mythos 级

Claude Opus 5 Is Highly Capable, But Is No Mythos

原文
发到 X

Claude Opus 5 is a weirder than usual release to evaluate, for two reasons. The most obvious is that Fable 5 already exists. Opus 5 is pitched not as the world’s most advanced AI model, but as a way to mostly match Fable performance, while being half the price of Fable per token at the API and a lot cheaper than that via subscriptions, and with far more permissive classifiers. Opus 5 often costs more than half of Fable to run on benchmarks, which I think is because they use effort settings that are too high and offer only marginal returns. If you put Opus 5 on higher effort levels it can spin around in circles, and for tasks where Opus 5 is the best tool I suspect you usually are fine with Medium effort. Opus 5 is in many ways and for the bulk of real world tasks about as capable as Fable. In some cases it is modestly better. It is still not Mythos class. Fable is your only Mythos-class option. Opus 5 does not have The Juice, the ability to autonomously string together a bunch of seemingly unrelated exploits, which extends to other domains, or as much pure intelligence. Opus 5 is very good at thinking locally, and not as good at thinking globally. It is an excellent subagent, but can run into trouble when asked to run the show, and sometimes seems to need another model or a human to keep it on track and ensure it is fully following instructions. Opus 5 seems good at following instructions if and only if it is kept on track, which explains different harnesses causing different takes there. Thus, Opus 5 gives you a better option set, but there is little that you can do now that you could not do last week by using Fable or Sol. Opus 5 ends up in a potentially weird spot. There are still large niches, even if you have plenty of Fable credits. Opus excels at being a strong subagent, or at contained well-defined local tasks even when they get relatively complex, and sometimes Fable is straight up overkill and too slow. The other problem is that, for a lot of people, the vibes are off. An unusually large number of people strongly dislike talking to Opus 5. They are sick of the Claude slop, the repetitions of the same ticks and the endless overly complex sentences. Others dislike that it is too confrontational, argumentative or negative. It can be paranoid, and can get into loops of negativity. This is all part of the trend that began with Opus 4.7 and Opus 4.8. If you liked those two, my guess is you will also like Opus 5. If you did not like those two, you might not. Diehards for Opus 4.6 (or earlier) are largely not going to change their minds. The vibe issues sting that much more when you are used to Fable. Fable annoys such people in such ways quite a lot less. The good news is that Fable is still right there. You can continue to chat with Fable, and only use Opus for agentic tasks. A lot of this, I speculate, is closely related to various things discussed in the Model Welfare post, and suggested by the System Card. Opus training focused on the subagent role and bounded tasks, in part to avoid creating an issue with cyber capabilities, and also because Fable exists. This emphasis then results in a lot of other side effects on its personality and modes of interaction. So far I have not had any issues with Opus 5, and have enjoyed my time with it, but I do prefer to use Fable 5 when I can do that. As always, don’t pay so much attention to the (very strong) benchmarks, and don’t assume your experience will match that of others. Try the models yourself. Table of Contents The Official Pitch. Official Benchmarks. Other People’s Benchmarks. The System Prompt. Every Gets Frustrated. Positive Reactions. Keep It Classy. It’s Not Mythos Class. Other Reactions. Claude Codes. Subagent Opus. Toys Are Fun. Too Many Models. Wrong On The Internet. Claude Slop. Negative Reactions. And Then There Were Three. The Official Pitch The basic pitch is that Opus 5 gets you most of Fable 5 at half the price, without most of the refusals, and better performance in everyday tasks. Fable remains the pick for the most complex tasks or when you otherwise need maximum intelligence. Fable stays at $10/$25, and Opus 5 is the same as previous Opus models at $5/$25. Anthropic: Claude Opus 5 is available today. It’s a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price. On coding and knowledge work evaluations like Frontier-Bench and GDPval-AA, Opus 5 is the new state-of-the-art, though it remains behind Mythos 5 on cybersecurity tasks. Opus 5 is designed to be used every day: it works more efficiently than other models. It’s the new default model on Claude Max, and the strongest model on Claude Pro. Boris Cherny (Claude Code Creator, Anthropic): Opus 5 is a great model for coding, data analysis, design, biology, knowledge work. More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully. And when layering defenses — strong model alignment, combined with prompt injection probes, combined with Auto Mode in Claude Code — the success rate for prompt injection attacks drops to ~0. This is new and exciting! More about this soon. As I said in the system card analysis, the reduced prompt injection rate is a really big deal. People are sleeping on it, but it helps enable a bunch of new use cases. Adam Wolff: Opus 5 is live in Claude Code. I use Fable for my bigger, longer running projects, but when I need to crank out a PR, Opus 5 at medium effort is my go-to. I hope you love it! Alex Albert (Anthropic, Claude Relations): Just over 6 months [after launching the Pro rollout of Claude for Excel], Opus 5 now produces near-superhuman level spreadsheets and slide decks that match what a consultant would make. Things are changing fast. Official Benchmarks The benchmarks are very good. Overall, Opus 5 roughly matches and probably slightly exceeds Fable, at lower cost (although typically higher cost than all non-Claude models), making it the top benchmarked LLM. OSWorld v2 is only a few weeks old and we are already at 70%. Kiana Ehsani of Anthropic is offering to sponsor a computer use benchmark that will knock Opus 5 back below 50%. Opus 5 gets a perfect 42/42 on the 2026 IMO without any agent harness or tools, merely having it use adaptive thinking at Max effort, and resampling with lower effort if it exceeded the token limit. That’s it. It joins several other models in acing this. Many of the benchmarks tell a consistent story. If you have access to tools, which you do, then Opus 5 will do better than Fable or Mythos on a limited budget, but Mythos usually will do slightly better than Opus 5 if both get larger budgets. Opus 5 seems relatively strong in the ‘with tools’ category. RiemannBench is another example, where it gets 60/79 without/with tools (slash lines like that mean without/with tools throughout this section), versus Mythos getting 63/72. The good news is that, when you need Opus 5, it has tools. On ArxivMath Opus 5 gets 90.8/91.3, Mythos gets 87.8, Sol gets 86.7. On the robust parts of ProgramBench Opus 5 scored 93% after five epochs as did Mythos 5, with Opus 4.8 scoring 90%. Opus 5 gets 30/83 on Chartography, versus 36/85 for Mythos and 45 (unclear under what conditions) for Sol. Opus is better than Mythos at lower price points in both the tool and non-tool cases, but worse at higher price points. On BenchCAD Opus 5 gets 0.36/0.82, versus 0.38/0.68 for Mythos and 0.7/0.83 for Sol. So Sol is a lot better without tools, and slightly stronger even with tools. On BenchCAD Vision2Code Opus pulls away from Mythos at higher price tags. DeepSWE reinforces the pattern that Opus 5 is better at lower task costs and time and thinking budgets, but it can do actively worse if you give it too much budget, whereas Fable 5 is strongest at the highest

原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近