跳到主内容
精选80Latent Space(RSS)模型发布/更新多源精选 ×12

Claude Opus 5发布:性能接近Fable,定价减半

[AINews] Claude Opus 5: Fable-level performance at Opus price (half Fable)

原文

In a rare Friday release, Opus 5 took the headlines today. Athrough most of its official benchmarks have it technically beating Fable, the official messaging still says it “comes close”. This mostly reflects the difficulty of Evals - today’s AIE track drop - not reflecting “big model smell” that Anthropic obviously knows Fable retains but can’t measure.Fortunately, independent evaluations of Opus confirm the outperformance:And the improved efficiency story, beyond just pricing, is also important… although it only just matches GPT 5.6 Sol:AI News for 7/23/2026-7/24/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!AI Twitter RecapTop Story: Claude Opus 5 model launchWhat happenedAnthropic’s Claude Opus 5 launch triggered a mix of benchmark scrutiny, strong anecdotal coding-agent praise, and renewed debate about frontier model evaluation.Multiple tweets explicitly discuss Claude Opus 5 as a newly launched model and compare it to other frontier systems on coding and general capability metrics, including Epoch’s ECI assessment, a FrontierCode anomaly discussion, and early user reactions from tool-use workflows like browser automation @abacaj, @abacaj.Epoch reported that Claude Opus 5 achieves an ECI of 159, “slightly below Fable 5’s value of 161,” while matching Fable 5 on SWE-ECI at 161 on software engineering benchmarks @EpochAIResearch.The ECI result immediately drew criticism from users who felt the score understated Opus 5’s practical improvements; one response called it “incredibly underrated,” noting it appears only 1 point better than Opus 4.8 despite seeming “much better at everything” in practice @scaling01. The same user argued for harder public benchmarks @scaling01.A separate thread highlighted an apparent benchmark irregularity: Opus 5 scored better on FrontierCode at medium effort than at higher effort, even though more effort improved performance on other evals @jerhadf. That suggests either task-specific search/effort tradeoffs or evaluation instability rather than monotonic gains from extra inference-time compute.Several technically literate users praised Opus 5’s coding performance. Mikhail Parakhin @MParakhin—said “Best-of-n rules” and reported a clear head-to-head win against Fable “for math and everything, really,” while wishing it were available in Codex.Arena promoted first impressions of Opus 5 and said leaderboard scores based on real-world use were coming soon @arena, indicating community evals were still catching up at posting time.Nous Research’s portal added access to the model, with a tweet saying users could directly use Opus 5 through Nous Portal and that a 20% discount applied to all models including Opus 5 @witcheer. This is distribution/availability rather than a capability claim.User anecdotes emphasized browser control / agentic tool use. One post said Opus 5 opened the browser and canceled a ChatGPT Pro subscription @abacaj, followed by “This thing can really drive a browser wow” @abacaj. These are isolated demos, not systematic evals, but they align with broader market interest in computer-use agents.Other early reactions were more memetic than technical, including “Opus 5 subway FPS result” @bijanbowen, “On Claude bro” @andrew_n_carr, and “They’re terrified of Anthropic” @teortaxesTex. These reflect sentiment but not evidence.Technical detailsEpoch Capabilities Index (ECI):Claude Opus 5 ECI = 159Fable 5 ECI = 161Claude Opus 5 SWE-ECI = 161, matching Fable 5 on software engineering @EpochAIResearchCommunity response noted the model appears only +1 ECI point vs Opus 4.8, which some readers considered too small relative to qualitative gains @scaling01, @scaling01.FrontierCode behavior: one evaluator noted medium-effort > high-effort on FrontierCode for Opus 5 despite the usual pattern of improvement with more effort elsewhere @jerhadf. The tweet does not provide raw numbers in this excerpt, but the central technical point is that increased effort was not uniformly beneficial.Anecdotal comparative claims:A clear head-to-head win vs Fable in one user’s testing, especially with best-of-n sampling @MParakhinMatching “mythos” in one ecosystem summary post, though without attached numbers @eliebakouchFacts vs opinionsMore factual / measurement-oriented claimsEpoch’s benchmark statement that Opus 5 scored 159 ECI and 161 SWE-ECI is the clearest empirical claim in the set @EpochAIResearch.Arena’s statement that first impressions are available and real-world leaderboard scores are forthcoming is factual but incomplete @arena.Nous Portal offering access to Opus 5 with a 20% discount is a product-availability fact @witcheer.Interpretations / opinions“ECI is underrated” and “we need harder public benchmarks” are opinions about benchmark validity and sensitivity @scaling01, @scaling01.“How to shake faith in any benchmark: show Anthropic doing meh on it” is rhetorical skepticism about benchmark discourse and community bias @teortaxesTex.“Best-of-n rules” and Opus being a “very clear winner” over Fable are informal practitioner judgments, useful but nonstandardized @MParakhin.“They’re terrified of Anthropic” and AGI-timeline speculation tied to Anthropic are pure opinion/speculation rather than launch evidence @teortaxesTex, @teortaxesTex.Different opinionsSupportive viewsThe strongest positive interpretation is that Opus 5 is materially stronger in real use than public aggregate benchmarks currently show, especially for coding and tool-use tasks.@MParakhin reports it beats Fable in his own testing and says best-of-n improves outcomes.@abacaj, @abacaj highlight effective browser automation, suggesting practical agentic competence.@bijanbowen calling the “subway FPS result” the best one yet implies visual/computer-use demo quality impressed viewers.@eliebakouch places Opus 5 among top closed-model releases and says it is “matching mythos,” framing it as a top-tier frontier entrant.Skeptical / critical viewsThe main criticism is not that Opus 5 is weak, but that benchmarking around it is unstable, underspecified, or misaligned with user impressions.@jerhadf points to a puzzling effort scaling inconsistency on FrontierCode.@scaling01 argues the ECI result seems too low relative to observed improvements and uses that to call for harder public benchmarks @scaling01.@teortaxesTex implies some benchmark trust is contingent and anthropic-specific results provoke benchmark criticism, i.e. social interpretation may be contaminating technical assessment.Neutral / analytic viewsEpoch’s framing is restrained: slightly below Fable overall, tied on SWE-specific capability @EpochAIResearch.Arena’s “first impressions now, real-world leaderboard later” is another neutral posture, effectively saying the community has not yet converged on a robust ranking @arena.ContextClaude-family models already had a reputation for strong coding performance, long-context utility, and relatively polished enterprise/product packaging, so Opus 5 entered a market where users were primed to test whether Anthropic could maintain or extend a coding lead.The launch lands amid a broader shift from static chat benchmarks toward agentic evaluations: browser use, tool invocation, parallel task execution, and software engineering loop completion. That is why even casual anecdotes like browser cancellation workflows gained attention—they map to a category of real-world competence that classic QA benchmarks miss.The benchmark friction around Opus 5 fits a wider ecosystem problem: aggregate capability scores often compress diverse behaviors into a single number. ECI and similar indices are useful for broad tracking, but one-number summaries can obscure:coding vs non-coding specializationinference-time compute/effort scaling behaviorbest-of-n gainstool-use reliabilityreal-world latency/cost tradeoffsThe FrontierCode “

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源
Opus 5 发布
gabriel原文
Opus 5 在 ARC-AGI-3 得分翻四倍
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)原文
Anthropic Opus 5 抗提示注入能力显著提升
Simon Willison 博客(RSS)原文
Opus 5自估41%概率为道德主体,较前代提升
AI Notkilleveryoneism Memes ⏸️原文
Anthropic 发布 Claude Opus 5,…
Anthropic Newsroom(web_list)原文
Anthropic发布Opus 5,更便宜且限制更少
TechCrunch AI(RSS)原文

相似阅读

另一事件,读法相近