跳到主内容
@wquguru
精选75Rohan Paul模型发布/更新

Kimi K3 在法律基准测试中大幅领先 Claude Fable 5

Kimi K3 nearly doubled its nearest rival Claude Fable 5, on a demanding benchmar…

原文
发到 X

Kimi K3 nearly doubled its nearest rival Claude Fable 5, on a demanding benchmark for autonomous legal work.

Kimi K3 at 26.7%, vs Claude Fable 5 at 14.2%.

The test covers 120 private assignments across 24 legal fields, including memos and deposition summaries.

Each model receives case files, works through them autonomously, then produces finished legal documents.

Every required rubric item must pass, so one missed detail fails the entire assignment. This strict grading explains why even the leader succeeds on only 26.7% of tasks.

But, overall, given 27 successful tasks per 100, it looks like we still have a long road ahead for completely unsupervised legal work by AI.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近