跳到主内容
@wquguru
精选75Hacker News Best(web_list)模型发布/更新多源精选 ×2

Qwen3.8 27B 开源模型发布,智能指数 52 登顶

Qwen3.8 27B 在 Artificial Analysis 得分 52

原文
发到 X

Artificial Analysis

人工智能分析

K

K

Alibaba

阿里巴巴

Open weights model

开放权重模型

Released August 2026

2026年8月发布

Qwen3.8 27B Intelligence, Performance & Price Analysis

Qwen3.8 27B 智能、性能与价格分析

CompareAPI Provider Benchmarks

比较API提供商基准测试

Model summary

模型摘要

Intelligence

智能

#1 / 135

#1 / 135

52

52

Artificial Analysis Intelligence Index

人工智能分析智能指数

4 out of 4 units for Intelligence.

智能指数4分满分中的4分。

Speed

速度

N/A

不适用

Output tokens per second

每秒输出令牌数

Unknown out of 4 units for Speed.

速度指数4分满分中的未知分数。

Cost

成本

In $0.00Out $0.00

输入$0.00 输出$0.00

N/A

不适用

Cost per Intelligence Index task

每项智能指数任务的成本

Unknown out of 4 units for Cost.

成本指数4分满分中的未知分数。

Verbosity

冗长性

#23 / 135

#23 / 135

160M

160M

Output tokens from Intelligence Index

智能指数输出的令牌数

4 out of 4 units for Verbosity.

冗长性指数4分满分中的4分。

Comparison Summary

比较摘要

Qwen3.8 27B is amongst the leading models in intelligence and well priced when comparing to other open weight models of similar size. The model supports text and image input, outputs text, and has a 256k tokens context window.

Qwen3.8 27B 在智能方面处于领先模型之列,与同类规模的其他开放权重模型相比,价格合理。该模型支持文本和图像输入,输出文本,并具有 256k 令牌的上下文窗口。

Qwen3.8 27B scores 52 on the Artificial Analysis Intelligence Index, placing it well above average among comparable models (median: 9). When evaluating the Intelligence Index, it generated 160M tokens, which is very verbose in comparison to the median of 43M.

Qwen3.8 27B 在人工分析智能指数上得分为 52,远高于同类模型的平均水平(中位数:9)。在评估智能指数时,它生成了 1.6 亿个令牌,与中位数 4300 万相比,非常冗长。

Pricing for Qwen3.8 27B is $0.00 per 1M input tokens (competitively priced, median: $0.04) and $0.00 per 1M output tokens (competitively priced, median: $0.15).

Qwen3.8 27B 的定价为每 100 万输入令牌 $0.00(价格具有竞争力,中位数:$0.04),每 100 万输出令牌 $0.00(价格具有竞争力,中位数:$0.15)。

Technical specifications

技术规格

ReasoningYesThis page shows the reasoning version of this model.A non-reasoning variant may also exist.
Input modalitySupports: text and image
Output modalitySupports: text
Context window256k~384 A4 pages of size 12 Arial font
Total parameters27B
LicenseApache 2.0
Model weightsHugging Face
推理是此页面显示此模型的推理版本。也可能存在非推理变体。
输入模态支持:文本和图像
输出模态支持:文本
上下文窗口256k~384 页 A4 纸,12 号 Arial 字体
总参数27B
许可证Apache 2.0
模型权重Hugging Face

135 models in this class

此类模型中的 135 个模型

Metrics are compared against models of the same class:

指标与同类模型进行比较:

  • Non-reasoning models → compared only with other non-reasoning models
  • Reasoning models → compared across both reasoning and non-reasoning
  • Open weights models → compared only with other open weights models of the same size class:
  • Tiny: ≤4B parameters
  • Small: 4B–40B parameters
  • Medium: 40B–150B parameters
  • Large: >150B parameters
  • Proprietary models → compared across proprietary and open weights models of the same price range, using a blended 3:1 input/output price ratio:
  • <$0.15 per 1M tokens
  • $0.15–$1 per 1M tokens
  • >$1 per 1M tokens
  • 非推理模型 → 仅与其他非推理模型比较
  • 推理模型 → 与推理和非推理模型进行比较
  • 开放权重模型 → 仅与相同大小类别的其他开放权重模型比较:
  • 微型:≤4B 参数
  • 小型:4B–40B 参数
  • 中型:40B–150B 参数
  • 大型:>150B 参数
  • 专有模型 → 与相同价格范围的专有和开放权重模型进行比较,使用混合 3:1 输入/输出价格比率:
  • <$0.15 每 100 万令牌
  • $0.15–$1 每 100 万令牌
  • >$1 每 100 万令牌

Highlights

亮点

Intelligence

智能

Artificial Analysis Intelligence Index · Higher is better

人工分析智能指数 · 越高越好

Speed

速度

Output tokens per second · Higher is better

每秒输出令牌数 · 越高越好

Cost per Task

每任务成本

Weighted average cost (USD) per Intelligence Index task · Lower is better

每项智能指数任务的加权平均成本(美元)· 越低越好

Prompt Options

提示选项

Intelligence

智能

Artificial Analysis Intelligence IndexUpdatedAgentic IndexUpdated

人工智能分析智能指数已更新代理指数已更新

Artificial Analysis Intelligence Index

人工智能分析智能指数

Artificial Analysis Intelligence Index v4.1.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR

人工智能分析智能指数 v4.1.1 包含 9 项评估:GDPval-AA v2、𝜏³-Banking、Terminal-Bench v2.1、SciCode、人类最后的考试、GPQA Diamond、CritPt、AA-Omniscience、AA-LCR

29 of 609 models

609 个模型中的 29 个

NEW

Add model from specific provider

从特定提供商添加模型

Reasoning models are indicated by a lightbulb icon

推理模型以灯泡图标标示

Artificial Analysis Intelligence Index

人工智能分析智能指数

Artificial Analysis Intelligence Index v4.1.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

人工智能分析智能指数 v4.1.1 包括:GDPval-AA v2、𝜏³-Banking、Terminal-Bench v2.1、SciCode、人类最后的考试、GPQA Diamond、CritPt、AA-Omniscience、AA-LCR。有关每项评估的详细分解及我们如何运行它们,请参阅智能指数方法论。

Open Weights / ProprietaryReasoning / Non-ReasoningText Only / Multimodal Inputs

开放权重 / 专有推理 / 非推理仅文本 / 多模态输入

Artificial Analysis Intelligence Index by Open Weights / Proprietary

按开放权重 / 专有划分的人工智能分析智能指数

Artificial Analysis Intelligence Index v4.1.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR

人工智能分析智能指数 v4.1.1 包含 9 项评估:GDPval-AA v2、𝜏³-Banking、Terminal-Bench v2.1、SciCode、人类最后的考试、GPQA Diamond、CritPt、AA-Omniscience、AA-LCR

29 of 609 models

609 个模型中的 29 个

NEW

Add model from specific provider

从特定提供商添加模型

ProprietaryOpen Weights (Commercial Use Restricted)Open Weights

专有开放权重(商业使用受限)开放权重

Reasoning models are indicated by a lightbulb icon

推理模型以灯泡图标标示

Artificial Analysis Intelligence Index

人工智能分析智能指数

Artificial Analysis Intelligence Index v4.1.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

人工智能分析智能指数 v4.1.1 包括:GDPval-AA v2、𝜏³-Banking、Terminal-Bench v2.1、SciCode、人类最后的考试、GPQA Diamond、CritPt、AA-Omniscience、AA-LCR。有关每项评估的详细分解及我们如何运行它们,请参阅智能指数方法论。

Open Weights

开放权重

Indicates whether the model weights are available. Models are labelled as 'Commercial Use Restricted' if the weights are available but commercial use is limited (typically requires obtaining a paid license).

指示模型权重是否可用。如果权重可用但商业使用受限(通常需要获得付费许可证),则模型被标记为“商业使用受限”。

Benchmarks

基准测试

Intelligence Evaluations

智能评估

Intelligence evaluations measured independently by Artificial Analysis · Higher is better

由Artificial Analysis独立测量的智能评估 · 越高越好

CodingTool UseLong ContextMultimodalInstruction FollowingFaithfulnessWritingUser InteractionBusinessFinanceLegalMedicalSee more

编码工具使用长上下文多模态指令遵循忠实度写作用户交互商业金融法律医疗查看更多

19 of 23 evaluations

23项评估中的19项

29 of 609 models

609个模型中的29个

NEW

Add model from specific provider

从特定提供商添加模型

GDPval-AA v2

GDPval-AA v2

Agentic real-world work tasks, (Elo-500)/2000

代理型真实世界工作任务,(Elo-500)/2000

𝜏³-BankingUpdated

𝜏³-Banking更新

Agentic tool use

代理型工具使用

Terminal-Bench v2.1

Terminal-Bench v2.1

Agentic coding & terminal use

代理型编码与终端使用

SciCode

SciCode

Coding

编码

Humanity's Last ExamUpdated

人类最后的考试更新

Reasoning & knowledge

推理与知识

GPQA Diamond

GPQA Diamond

Scientific reasoning

科学推理

CritPt

CritPt

Physics reasoning

物理推理

AA-Omniscience AccuracyUpdated

AA-Omniscience准确性更新

Knowledge

知识

AA-Omniscience Non-Hallucination RateUpdated

AA-全知非幻觉率(更新)

1 - hallucination rate

1 - 幻觉率

AA-LCRUpdated

AA-LCR(更新)

Long context reasoning

长上下文推理

AA-Briefcase

AA-公文包

Agentic knowledge work, Elo

代理式知识工作,Elo

AutomationBench-AA

AutomationBench-AA

Agentic SaaS workflows

代理式SaaS工作流

Harvey LAB-AA

Harvey LAB-AA

Legal agentic work, criterion pass rate

法律代理工作,标准通过率

EnterpriseOps-Gym-AA

EnterpriseOps-Gym-AA

Agentic business operations

代理式业务运营

AA-AnalystAgentNew

AA-分析师代理(新)

Quantitative analysis on spreadsheets & documents

电子表格与文档的定量分析

IFBench

IFBench

Instruction following

指令遵循

APEX-Agents-AA

APEX-Agents-AA

Long-horizon agentic tasks

长周期代理任务

ITBench-AA

ITBench-AA

Kubernetes incident root-cause analysis

Kubernetes事件根因分析

MMMU-Pro

MMMU-Pro

Visual reasoning

视觉推理

Reasoning models are indicated by a lightbulb icon

推理模型以灯泡图标标示

Intelligence Evaluation Relevance

智能评估相关性

While model intelligence generally translates across use cases, specific evaluations may be more relevant for certain use cases.

虽然模型智能通常在不同用例中具有通用性,但特定评估可能对某些用例更为相关。

Artificial Analysis Intelligence Index

人工分析智能指数

Artificial Analysis Intelligence Index v4.1.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

人工分析智能指数 v4.1.1 包括:GDPval-AA v2、𝜏³-Banking、Terminal-Bench v2.1、SciCode、人类最后的考试、GPQA Diamond、CritPt、AA-Omniscience、AA-LCR。有关每项评估及其运行方式的详细分解,请参阅智能指数方法论。

AA-Omniscience

AA-Omniscience

AA-Omniscience IndexAA-Omniscience AccuracyAA-Omniscience Hallucination Rate

AA-Omniscience 指数AA-Omniscience 准确率AA-Omniscience 幻觉率

AA-Omniscience Index

AA-Omniscience 指数

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.

AA-Omniscience 指数(越高越好)衡量知识可靠性和幻觉。它奖励正确答案,惩罚幻觉,对拒绝回答不设惩罚。分数范围从 -100 到 100,其中 0 表示正确与错误答案数量相当,负分表示错误多于正确。

29 of 480 models

480 个模型中的 29 个

NEW

Add model from specific provider

从特定提供商添加模型

Reasoning models are indicated by a lightbulb icon

推理模型以灯泡图标标示

AA-Omniscience Index

AA-Omniscience 指数

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.

AA-Omniscience 指数(越高越好)衡量知识可靠性和幻觉。它奖励正确答案,惩罚幻觉,对拒绝回答不设惩罚。分数范围从 -100 到 100,其中 0 表示正确与错误答案数量相当,负分表示错误多于正确。

Openness Index

开放性指数

Openness IndexOpenness Index ComponentsOpenness vs. Intelligence

开放性指数开放性指数组成部分开放性 vs. 智能

Artificial Analysis Openness Index: Score

人工分析开放性指数:得分

Openness Index assesses model openness on a 0 to 100 normalized scale (higher is more open)

开放性指数在 0 到 100 的归一化尺度上评估模型开放性(越高表示越开放)

20 of 306 models

306 个模型中的 20 个

NEW

Add model from specific provider

从特定提供商添加模型

Reasoning models are indicated by a lightbulb icon

推理模型以灯泡图标标示

Intelligence Index Comparisons

智能指数比较

Intelligence Index vs. Cost per TaskIntelligence Index vs. Time per TaskIntelligence Index vs. Output SpeedIntelligence Index vs. End-to-End Response Time

智能指数 vs. 每任务成本智能指数 vs. 每任务时间智能指数 vs. 输出速度智能指数 vs. 端到端响应时间

Intelligence Index vs. Cost per Intelligence Index Task

智能指数 vs. 每项智能指数任务的成本

Artificial Analysis Intelligence Index · Weighted average cost (USD) per Artificial Analysis Intelligence Index task

人工智能分析智能指数 · 每项人工智能分析智能指数任务的加权平均成本(美元)

29 of 609 models

609个模型中的29个

NEW

Most attractive quadrant

最具吸引力的象限

Pareto line

帕累托线

GoogleAnthropicZ AIDeepSeekSpaceXAIKimiNVIDIAMetaOpenAIAlibabaMistral

谷歌Anthropic Z AI DeepSeek SpaceX AI Kimi NVIDIA Meta OpenAI 阿里巴巴 Mistral

Reasoning models are indicated by a lightbulb icon

推理模型以灯泡图标标示

Cost per Intelligence Index Task

每项智能指数任务的成本

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近