跳到主内容
@wquguru
精选88Rohan Paul模型发布/更新多源精选 ×8

Gemini 3.8 Flash 发布,基准测试超越 Claude Opus 5

Gemini 3.8 Flash is out.

原文
发到 X
推荐理由

Gemini 3.8 Flash 直接对标 Claude Opus 5 并在多项关键基准取得领先,对评估当前顶级模型能力有重要参考价值,建议关注其实际部署表现。

Gemini 3.8 Flash is out.

beats Claude Opus 5 on some really important benchmarks.

  • 54.9% on HLE-Verified vs 54.4% for Claude Opus 5
  • Terminal-Bench 2.1, that test tests whether an agent can actually operate a terminal and successfully finish difficult coding, security, ML, data-science, and systems tasks. 89.4% is essentially a very high task-completion rate under that evaluation setup.
  • Harvey's Legal Agent Benchmark, Gemini 3.8 Flash scores 10.0% vs Opus 5's 6.7%. The number looks low because this uses an extremely strict all-pass rule: a legal workflow gets credit only when every required criterion passes, including facts, conclusions, citations, structure, and analysis, across complex file-based legal work.

Google has not raised the price per token. On harder tasks, Gemini 3.8 Flash may think longer and make more tool calls, so it uses more tokens and the total cost of completing that task can rise even though the token price stays the same.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近