Kimi K3 智能体知识工作排名第二,单任务成本 10.57 美元
Kimi K3 now ranks 2nd in agentic knowledge work, but each task costs $10.57 to r…
Kimi K3 now ranks 2nd in agentic knowledge work, but each task costs $10.57 to run, roughly 10x more than K2.6, and above Opus.
Very near to Fable-5 on AA-Briefcase, a private agentic work benchmark.
The AA-Briefcase test hands models messy real jobs, then grades the spreadsheets, slides, and mockups they build. So this benchmark measures finished deliverables across long tasks rather than grading one isolated answer.
Kimi K3 scored an Elo of 1543, Claude Fable 5 at 1574. That beats GPT-5.6 Sol at 1501, Claude Sonnet 5 at 1388, and Claude Opus 4.8 at 1347.
The jump from the last version is where the real success story. Kimi K2.6 sat at 816, so this is a +727 leap in one generation. On raw correctness it passed 51% of the rubric, behind only Fable 5 at 56%. Its analytical Elo of 1754 nearly matches Fable 5 at 1744.
Now the catch, which is the price. Each task runs $10.57, roughly 10x more than K2.6, and above Opus. That comes from 83 turns per job and heavy output, stretching each task to 56 minutes. i.e. Kimi K3's long runs generate more output tokens, which carry Kimi’s highest per-token price.
The model also revisits tasks repeatedly, so spending grows across every added reasoning cycle.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力