跳到主内容
@wquguru
精选75Hacker News Best(web_list)模型发布/更新

Qwen3.8 Max登顶Agentic Index,成综合最强模型

Qwen3.8 Max now ranked as the best overall model by agentic index

原文
发到 X

Artificial Analysis

K

Independent analysis of AI

Understand the AI landscape to choose the best model and provider for your use case

Update

Intelligence Index v4.1.1

Intelligence Index v4.1.1 moves 𝜏³-Banking to v1.0.1 and upgrades the grader for HLE, AA-LCR, and AA-Omniscience to GPT-5.6 Luna (medium)

Launch

Endpoint Accuracy Index

Measuring whether provider endpoints serve the same model quality as the reference

Highlights

Intelligence

Artificial Analysis Intelligence Index · Higher is better

Speed

Output tokens per second · Higher is better

Cost per Task

Weighted average cost (USD) per Intelligence Index task · Lower is better

Personalized model recommender

Get personalized recommendations based on your priorities for intelligence, speed, and cost

Explore agents for general work, coding, customer support, and more

Compare AI agents across capabilities, pricing, and platform support

Explore premium plans

Access expanded benchmark data, custom visualizations, industry reports, and more

Changelog

New article published · 6 Aug

Launching v4.1.1 of the Artificial Analysis Intelligence Index

Methodology updated · 6 Aug

Artificial Analysis Intelligence Index v4.1.1

New language model evaluation · 6 Aug

Ling 3.0 Tiny

New article published · 5 Aug

Muse Spark 1.2

New language model evaluation · 5 Aug

Qwen3.8 Max

New language model evaluation · 5 Aug

Ling-3.0-flash

New language model evaluation · 5 Aug

Muse Spark 1.2 (xhigh)

New article published · 4 Aug

Launching the Endpoint Accuracy Index: Same Model, Different Accuracy

New language model evaluation · 3 Aug

G9v3-39A5B

New article published · 31 Jul

DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index, 10 points above previous DeepSeek V4 Flash

New language model evaluation · 31 Jul

Celeris-1

New language model evaluation · 31 Jul

DeepSeek V4 Flash 0731 (Reasoning, Max Effort)

New article published · 30 Jul

Inkling Small lands within a point of Inkling on the Artificial Analysis Intelligence Index with less than a third of the parameters

Methodology updated · 30 Jul

We have updated our Cost per Task methodology, resulting in slight absolute increases in cost estimates but with minimal impact on relative positioning.

New language model evaluation · 30 Jul

Kimi K3 (low)

New language model evaluation · 30 Jul

Inkling Small

New article published · 29 Jul

Agnes AI releases Agnes 2.5 Pro Alpha

New article published · 24 Jul

Claude Opus 5: the new leader in agentic knowledge work

New article published · 24 Jul

Opus 5: Fable 5 level intelligence at a lower cost per task

New language model evaluation · 24 Jul

Claude Opus 5 (Adaptive Reasoning, Low Effort)See more

Intelligence

Intelligence of leading AI models based on our independent evaluations

Artificial Analysis Intelligence IndexUpdatedAgentic IndexUpdated

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.1.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR

26 of 595 models

NEW

Add model from specific provider

Estimate (independent evaluation forthcoming)

Reasoning models are indicated by a lightbulb icon

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.1.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Open Weights / ProprietaryReasoning / Non-ReasoningText Only / Multimodal InputsBy Country

Artificial Analysis Intelligence Index by Open Weights / Proprietary

Artificial Analysis Intelligence Index v4.1.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR

26 of 595 models

NEW

Add model from specific provider

Estimate (independent evaluation forthcoming)

ProprietaryOpen Weights (Commercial Use Restricted)Open Weights

Reasoning models are indicated by a lightbulb icon

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.1.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Open Weights

Indicates whether the model weights are available. Models are labelled as 'Commercial Use Restricted' if the weights are available but commercial use is limited (typically requires obtaining a paid license).

Cost per TaskTime per TaskOutput Tokens per Task

Cost per Intelligence Index Task

Weighted average cost (USD) per Artificial Analysis Intelligence Index task, segmented by token type. Lower is better

26 of 595 models

NEW

AnswerReasoningCache WriteCache HitInput

Reasoning models are indicated by a lightbulb icon

Cost per Intelligence Index Task

Weighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight.

Intelligence Index vs. Cost per TaskIntelligence Index vs. Time per TaskIntelligence Index vs. Output Tokens per Task

Intelligence Index vs. Cost per Intelligence Index Task

Artificial Analysis Intelligence Index · Weighted average cost (USD) per Artificial Analysis Intelligence Index task

26 of 595 models

NEW

Most attractive quadrant

Pareto line

XiaomiMetaGoogleAnthropicMiniMaxOpenAINVIDIAAlibabaSpaceXAIMistralZ AIKimiDeepSeek

Reasoning models are indicated by a lightbulb icon

Cost per Intelligence Index Task

Weighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight.

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.1.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Frontier Language Model Intelligence, Over Time

Artificial Analysis Intelligence Index v4.1.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR

14 of 57 model creators

NEW

AnthropicOpenAIKimiAlibabaMetaSpaceXAIZ AIDeepSeekGoogleMiniMaxXiaomiThinking MachinesMistralCohere

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.1.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Coding Agent Index

Performance, cost, and execution time for leading coding agents on end-to-end software engineering tasks

IndexCostExecution Time

Artificial Analysis Coding Agent Index

Composite average pass@1 across DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA · Higher is better

Color by

ModelAgent

15 of 52 models

NEW

Image & Video

Top models from our Image Arena and Video Arena leaderboards, with 95% confidence intervals

Text to ImageImage EditingText to VideoImage to VideoVideo Editing

Text to Image Leaderboard

Elo scores from blind preference votes in our Image Arena. See the full leaderboard here.

15 of 150 models

NEW

Speech

Top models from our Text to Speech Arena, Speech to Text and Speech to Speech evaluations

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源
阿里Qwen3.8登顶Agentic能力榜
量子位(RSS)原文

相似阅读

另一事件,读法相近