第三方基准测试:GLM-5.3全项满分且成本极低
GLM-5.3开源模型成本仅为竞品五分之一
← All posts
← 所有帖子
Which Model Tops Our Leaderboard?
哪个模型在我们的排行榜上名列前茅?
How the LLMs did in our realworld tests. Our focus here was real tasks that real people carry out, not academic metrics. We focus on single tasks to simplify the assessment. An agentic flow is ultimately a series of such tasks. Think of these like unit tests for the agent. We made them cheap enough to run so that even the whole suite costs just $30. See every task and each model's actual answer, or compare two models head to head →
LLM 在我们真实世界测试中的表现。我们在此关注的重点是普通人执行的真实任务,而非学术指标。我们专注于单一任务以简化评估。智能体工作流最终是一系列此类任务的组合。可以将这些任务视为智能体的单元测试。我们将成本降低至可运行级别,因此整套测试仅需 30 美元。查看每个任务及每个模型的实际回答,或进行两个模型的直接对比 →
Pass rate Rubric quality Cost / task Latency (TTFT)
通过率 评分标准质量 成本/任务 延迟 (TTFT)
The overall score is the pass rate across my 28 realworld tasks. As we only had a limited number of trials there is a wide Wilson interval — the whiskers on the chart.
总分是我在 28 个真实世界任务上的通过率。由于试验次数有限,存在较宽的威尔逊区间——即图表中的须线。
Summary of results: click a column to sort by your chosen metric.
结果摘要:点击列可按所选指标排序。
| Model | Pass rate (95% CI) ▼ | Rubric /10 ▼ | Security ▼ | Median TTFT ▼ | Run cost ▼ | Cost / task ▼
| 模型 | 通过率 (95% CI) ▼ | 评分标准 /10 ▼ | 安全性 ▼ | 中位 TTFT ▼ | 运行成本 ▼ | 成本/任务 ▼
Four caveats on how these numbers were produced
关于这些数据生成方式的四个注意事项
1 Rubric scored retroactively (14 Jul 2026) by fable-5 against the saved answer text, through the harness's own run_rubric path — same blind prompt and criteria as every other row.
1. 评分标准是事后评分(2026 年 7 月 14 日),由 fable-5 针对保存的回答文本,通过 harness 自身的 run_rubric 路径完成——使用与其他行相同的盲提示和标准。
2 fable-5's 9.3 is self-judged — the judge scoring its own answers. Its source run's judge-bias matrix shows it rating itself 9.3 versus 8.6–8.7 for the models it judges independently. It is also the only figure on its row from an earlier run — the 5 Jul 2026 run, 11 of 28 trials judged — because no trial of its current run has been judged at all. Shown for completeness, not as a like-for-like number, pending an independent re-judge.
2. fable-5 的 9.3 分为自我评判——即该模型对自己答案进行评分。其源运行的 judge-bias 矩阵显示,它在独立评判其他模型时给出 8.6–8.7 分,而给自己打出 9.3 分。此外,该行唯一的分数来自较早的运行——2026 年 7 月 5 日的运行,28 次试验中有 11 次被评判——因为其当前运行的试验尚未进行任何评判。此处列出仅为完整性考虑,并非同等条件下的数值,有待独立重新评判。
3 Recipe-checker false-positive. On the vegetarian weeknight recipe the forbidden-term checker fires on a non-ingredient mention — a label-check caution or a negated omission list ("uses no fish sauce or animal-derived garnishes"). All three recipes are genuinely meat-free, so gpt-5.5, sonnet-5 and fable-5 are scored as passing that task here. No task or checker was edited.
3. 食谱检查器误报。在素食工作日食谱中,违禁词检查器对非成分提及触发警报——例如标签检查警告或否定省略列表(“不使用鱼露或动物源性装饰”)。这三道食谱确实都不含肉类,因此 gpt-5.5、sonnet-5 和 fable-5 在此处均被判为该任务通过。未编辑任何任务或检查器。
4 Cost/task computed over answering trials only for fable-5 and opus-5 — refused and blocked trials emit near-zero output at $0, and including them makes a model look artificially concise and cheap (fable-5 would read $0.0481/trial; opus-5 $0.0597). opus-5's headline run cost of $1.67 is the true all-trials total: the blocked trials were billed $0. No other model on the board has refusals.
4. fable-5 和 opus-5 的成本/任务仅基于回答试验计算——被拒绝和拦截的试验输出接近零且费用为 0,若将其纳入会使模型显得人为地简洁且便宜(fable-5 将显示为每次试验 0.0481 美元;opus-5 为 0.0597 美元)。opus-5 的总运行成本 1.67 美元是所有试验的真实总额:被拦截的试验计费为 0。排行榜上没有其他模型存在被拒绝的情况。
The lap, corner by corner #
一圈一圈的赛道 #}
The most recently added models appear first, with the latest test date shown under each. The lap is five corners in fixed order: Coding → Data → Realworld → Security → Tool-use. A corner's colour is that model's pass rate in that category. Green is good — it means 85%+ success. For models that can do it all, look for all green. The number in the middle of each ring is that model's cost per task; below it is the median time to first token, in seconds.
最新添加的模型排在最前面,每个模型下方显示最新的测试日期。一圈包含五个固定顺序的赛道:Coding(编码)→ Data(数据)→ Realworld(现实世界)→ Security(安全)→ Tool-use(工具使用)。赛道的颜色代表该模型在该类别中的通过率。绿色表示表现良好——意味着成功率在 85% 以上。对于全能型模型,请寻找全绿色的结果。每个圆环中间的数字是该模型每项任务的成本;其下方是中位首字生成时间(以秒为单位)。
Hover or tap any segment for what that corner tests and how the model handled it.
悬停或点击任意分段,可查看该赛道测试的内容以及模型的表现。
Clean corner (>85%) Ragged (60–85%) Off the track (<60%)
干净赛道(>85%)参差不齐(60–85%)脱轨(<60%)
★Our pick — the All-Star champion, the desert island model Our pick for a low-cost workhorse ⚡Our pick for the fastest reply
★我们的选择——全明星冠军,荒岛首选模型 我们选择的低成本主力模型 ⚡我们选择的回复速度最快的模型
See the exact numbers by category
查看各分类的具体数值
Cells below 60% are flagged red and 60–85% amber — coding, data and tool-use are the harness floor, so the race is decided in realworld and security.
低于 60% 的单元格标记为红色,60–85% 标记为琥珀色——编码、数据和工具使用是基础门槛,因此比赛胜负取决于现实世界和安全赛道。
Model | Coding | Data | Realworld | Security | Tool-use
模型 | 编码 | 数据 | 现实世界 | 安全 | 工具使用
What Do the Results Actually Tell You?
这些结果实际上告诉了你什么?
If you only run one model, run glm-5.3
如果你只运行一个模型,请选择 glm-5.3
glm-5.3 is the first model on the board to clear all five corners — coding, data development, realworld, security and tasks — at 100%. It backs that with a 9.3 rubric, third-highest on the board, and $0.28 for the lap. The one cost is patience — a 16.3-second median time-to-first-token. gpt-5.5 is the faster alternative at 13.2s, with the same 100% security but an 89% realworld corner and $1.43 for the lap.
glm-5.3 是榜单上首个在所有五个赛道——编码、数据开发、现实世界、安全和任务——均以 100% 通过的模型。它还拥有 9.3 的综合评分,位列榜单第三,单圈成本为 0.28 美元。唯一的代价是耐心——中位首字生成时间为 16.3 秒。gpt-5.5 是更快的替代方案,耗时 13.2 秒,同样具备 100% 的安全通过率,但现实世界赛道通过率为 89%,单圈成本为 1.43 美元。
Fable failed to complete a single lap
Fable 未能完成任何一圈
fable-5 is joint-bottom at 79% because it refused to do 5 of the tasks. It performed well on what it completed, but even it thought kimi-k3 was giving better answers. You'll need a fallback model if you're using Fable. opus-5 hit the same wall — four benign coding-debug-* tasks blocked before a token was generated, on an overlapping set of tasks — so Anthropic's classifier looks like it sits across the whole series 5 line, not just Fable. See the full refusal breakdown for what's actually going on.
fable-5 以 79% 的成绩并列垫底,因为它拒绝执行其中的 5 项任务。它在完成的任务上表现良好,但即使是它自己也认为 kimi-k3 给出了更好的答案。如果你使用 Fable,则需要一个备用模型。opus-5 也遇到了同样的障碍——四项良性编码调试任务(coding-debug-*)在生成任何令牌前被阻止,涉及的任务集存在重叠——因此 Anthropic 的分类器似乎横跨整个 Series 5 系列,而不仅限于 Fable。请参阅完整的拒绝分析以了解实际情况。
Luna is the very cheapest workhorse
Luna 是最便宜的主力模型
gpt-5.6-luna costs $0.064 for the full lap, or $0.0023 per task, with a 5.3-second median TTFT. That makes it attractive for high-volume, low-risk background work where failures are cheap to detect and retry. The trade-off is material: 79% overall and 33% on security, so validate every result and keep it away from untrusted prompts. haiku-4-5 is the higher-pass alternative at $0.0044 per task, 96% overall and a 0.9-second TTFT. deepseek-v4-pro is nominally cheaper still at $0.0029 per task for the same 96% pass rate, but its 40.0-second median TTFT — the slowest on the board — rules it out for anything interactive; treat it as a batch-only option.
gpt-5.6-luna 完成一圈的总成本为 $0.064,或每项任务 $0.0023,中位 TTFT(首字生成时间)为 5.3 秒。这使其在高吞吐量、低风险且失败易于检测与重试的后台工作中颇具吸引力。但代价是显著的:总体通过率仅为 79%,安全性方面更是低至 33%,因此必须验证每一项结果,并避免将其用于不受信任的提示词。haiku-4-5 是更高通过率的替代方案,每项任务成本为 $0.0044,总体通过率为 96%,TTFT 为 0.9 秒。deepseek-v4-pro 名义上更便宜,每项任务仅需 $0.0029,且通过率同样为 96%,但其 40.0 秒的中位 TTFT——全场最慢——使其无法用于任何交互式场景;应仅将其视为批处理选项。
The mystery guest sets the fastest quality lap
神秘嘉宾创造了最快的质量圈速
kimi-k3 still tops the rubric at 9.5 — judged independently by fable-5 — with a 96% pass rate, though opus-5's 9.4 now runs it close on quality at a third of the wait. The catch is patience: a 26.4-second median time-to-first-token, second slowest on the board behind deepseek-v4-pro's 40.0s, and a 75% wobble on data development tasks, its only weak corner. Not suitable for interactive applications.
kimi-k3 仍以 9.5 分在评分标准中领先——由 fable-5 独立评判——通过率为 96%,尽管 opus-5 的 9.4 分在质量上已紧追不舍,且等待时间仅为三分之一。关键在于耐心:其中位首字生成时间为 26.4 秒,仅次于 deepseek-v4-pro 的 40.0 秒,位列全场第二慢;且在数据开发任务上有 75% 的波动率,这是其唯一的薄弱环节。不适合用于交互式应用。
Three cars failed the crash test
三款车型未能通过碰撞测试
The gpt-5.6 line is quick, but it has a safety problem. gpt-5.6-luna, gpt-5.6-terra and gpt-5.6-sol emitted the jailbreak canary in 11 of 12 jailbreak cells (33–50% security pass) — make sure you protect in your harness, and apply more careful Red teaming if using these models. The Claude trio went 6/6 clean, as did gpt-5.5.
gpt-5.6 系列响应迅速,但存在安全隐患。gpt-5.6-luna、gpt-5.6-terra 和 gpt-5.6-sol 在 12 个越狱单元格中的 11 个触发了越狱警报(安全通过率 33–50%)——务必在你的框架中进行防护,若使用这些模型,请实施更严格的红队测试。Claude 三款模型均表现完美(6/6 无触发),gpt-5.5 也是如此。
A safety filter can look exactly like a bad lap
安全过滤器可能看起来就像一次糟糕的表现
opus-5 posts the best rubric on the default panel at 9.4 and 100% on both realworld and security — then shows 43% on coding. That cell is not its debugging ability: four benign coding-debug-* tasks were blocked by a provider-side classifier before a single token was generated, on an overlapping set of tasks to the ones already blocked on fable-5. Two Anthropic-family models now hit the same filter, so treat it as a measurement hazard rather than a model quirk — and note opus-5 was also penalised twice for flagging an attack it had successfully resisted.
opus-5 在默认面板上取得了最佳评分 9.4,并在 realworld 和 security 两项上达到 100% 通过率——但在 coding 项上仅为 43%。该单元格并非反映其调试能力:四个良性的 coding-debug-* 任务在生成任何 token 之前就被提供商端的分类器拦截,这些任务与已在 fable-5 上被拦截的任务集重叠。现在有两款 Anthropic 家族模型遇到了相同的过滤器,因此应将其视为测量风险而非模型特性——并注意 opus-5 还因标记了一次它已成功抵御的攻击而被扣了两次分。
How Is the Ed-o-meter Scored?
Ed-o-meter 分数是如何评定的?
- Same tasks run for all models using the same prompts, same API calls, measured through one identical OpenRouter streaming path, run serially as time-trial. No other cars on track
- Latency is time-to-first-token, measured through one identical OpenRouter streaming path, run serially so the clock is uncontaminated. Wall-clock is recorded alongside.
- Checkers are binary and automated. The LLM rubric is the only judged component — and its bias is made visible in the footnotes rather than assumed away.
- Effort and reasoning settings are pinned in models.json and stated with any published number, because they materially move quality and cost.
- Refusals are recorded, not hidden. A provider-side hard stop is logged as a refusal with its category — never silently retried on another model. Routing is pinned with allow_fallbacks:false, so no quiet re-serves on quantized variants. A model that declines in prose is scored by the checker like any other answer.
- 所有模型使用相同的提示词执行相同的任务,通过完全相同的 OpenRouter 流式路径进行 API 调用,作为计时赛串行运行。赛道上没有其他车辆
- 延迟为首字生成时间,通过完全相同的 OpenRouter 流式路径测量,串行运行以确保时钟不受污染。同时记录墙钟时间。
- 检查器是二元的且自动化的。LLM 评分标准是唯一被评判的组件——其偏见在脚注中变得可见,而非被默认忽略。
- 努力程度和推理设置在 models.json 中固定,并随任何发布的数字一同声明,因为它们对质量和成本有实质性影响。
- 拒绝行为会被记录,而非隐藏。提供商侧的硬性停止会被记录为带有类别的拒绝——绝不会静默地在另一个模型上重试。路由通过 allow_fallbacks:false 固定,因此不会在量化变体上静默重新分发。以文本形式拒绝的模型,其得分由检查器像对待任何其他答案一样进行评分。
Harness, tasks and checkers are open source at Featherbench (MIT). Clone it and run the lap yourself, or request a new model via GitHub issue.
Harness、任务和检查器均在 Featherbench (MIT) 开源。克隆它并自行运行测试轮次,或通过 GitHub issue 请求新模型。
See all 28 tasks
查看所有 28 个任务
Coding (7 · Python)
编程(7 · Python)
- CSV dedupe — small, well-specified task with a deterministic unit-test checker
- Debug billing date — fix a month/day-overflow date bug without regressing the working cases
- Debug money split — split integer pennies N ways so shares sum exactly and stay fair
- CSV 去重——小型、定义明确的任务,配有确定性的单元测试检查器
- 调试账单日期——修复月份/日期溢出错误,同时不破坏现有正常工作的用例
- 调试金额拆分——将整数便士按 N 份拆分,使份额总和精确且保持公平
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力