跳到主内容
精选88Latent Space(RSS)模型发布/更新多源精选 ×11

Kimi K3 发布:2.8T参数开源模型,性能对标Opus 4.8

[AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing

原文
推荐理由

这是目前最大的开源模型,性能接近闭源旗舰,做 Agent 和长上下文应用的团队值得关注其权重开放后的实际表现。

Z.ai GLM has been getting a bit too much love recently, so it’s time for Kimi K3 to fight back! It’s hard to put the scale of today’s open model release in perspective, so thankfully Moonshot AI did it for us:Their vibe reel was entirely edited by Kimi K3 and worth a watch:You can read SimonW and Arena for standard takes and rankings, none of which will be particularly unexpected given the large size of the model, but this pic best summarizes the K2.5 to K3 jump:AI News for 7/15/2026-7/16/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!AI Twitter RecapMoonshot AI launched Kimi K3 as a frontier-class open-weights model, with official claims that place it near top closed models and above prior open competitors.Moonshot officially introduced Kimi K3 as “Open Frontier Intelligence” with 2.8T total parameters, 1M-token context, native multimodal input, Kimi Delta Attention (KDA), and Attention Residuals, and said the model is live on Kimi.com, Kimi Work, Kimi Code, and API, with open weights promised by July 27, 2026 @Kimi_MoonshotMoonshot also highlighted product positioning around long-horizon agentic coding and self-evolving workflows, plus “vision in the loop” coding/game-building workflows that iterate between code and screenshots @Kimi_MoonshotBefore the formal announcement, multiple accounts circulated leaked or app-sourced details that K3 was 2.8T params, calling it the largest open-weight model ever if weights ship as promised @scaling01, @scaling01, @eliebakouchThe official Kimi blog went live later and was widely shared as the primary technical source @Jianlin_S, @scaling01, @Yulun_DuMoonshot’s own phrasing acknowledged a limitation: despite being highly competitive overall, K3 still has a “noticeable gap in user experience” versus Claude Fable 5 and GPT-5.6 Sol @scaling01Arena announced that Kimi K3 entered Agent Arena, plus Text, Vision, Document, and Frontend Code Arena, with community evaluations to follow @arenaArena then reported a major early result: Kimi K3 became #1 in Frontend Code Arena with 1679 points, surpassing Claude Fable 5 and jumping from #18 (K2.6) to #1, ranking #1 in 6 of 7 frontend domains and #2 in Gaming @arenaArena later added that K3 has a 76% pairwise win rate in Frontend Code Arena, versus 63% for Fable 5 and 58% for GPT-5.6 Sol @arenaIn Text Arena, K3 landed at #9 with 1486 points, a jump from #38, with top-10 placements in creative writing, coding, and instruction following, and #1 in several occupation slices @arenaArtificial Analysis published an independent evaluation placing K3 at 57 on the AA Intelligence Index, calling it comparable to Opus 4.8 and GPT-5.5, but still behind Fable 5 and GPT-5.6 Sol overall @ArtificialAnlysAA also reported K3 at 1668 Elo on GDPval v2, 53% / #1 on AutomationBench-AA, and 1547 Elo on AA-Briefcase, with cost per task of $0.94, about 21% fewer output tokens than K2.6 across the full Intelligence Index run @ArtificialAnlysThe launch immediately triggered strong reaction from engineers and model-watchers who framed K3 as an open-model milestone comparable to earlier DeepSeek moments @kimmonismus, @nrehiew_, @eliebakouchTechnical detailsArchitecture and systems detailsOfficial specs: 2.8T total parameters, 1M context, native multimodal input (text + images), text output, open weights by July 27 @Kimi_Moonshot, @ArtificialAnlysK3 uses Kimi Delta Attention (KDA), which Moonshot says enables up to 6.3x faster decoding in million-token contexts @Kimi_MoonshotIt also uses Attention Residuals (AttnRes), claimed to deliver ~25% higher training efficiency at <2% additional cost @Kimi_MoonshotCommunity readers of the blog highlighted additional architecture details: LatentMoE / Stable LatentMoE, 16 activated experts out of 896, implying an activation ratio under 2% @nrehiew_, @eliebakouchMore community-extracted details from the blog/report discussion: per-head Muon, QB load balancing / quantile load balancing, and a new activation function called SiTU (Sigmoid Tanh Unit) @eliebakouchOne engineer noted the architecture as notable for combining KDA + LatentMoE + AttnRes while scaling more than 2x over prior Kimi models @teortaxesTexKDA had a long incubation cycle: design reportedly started in Jan 2025 and took ~1.5 years to reach frontier scale @zxytimInference and servingK3 pricing was reported as $3 / 1M input tokens and $15 / 1M output tokens, with cached input discounted 90% to $0.30 / 1M @scaling01, @ArtificialAnlysSeveral posters compared that pricing to Sonnet 5, with some noting Sonnet was temporarily cheaper until end of August, after which prices align more closely @kimmonismusA blended estimate at 80% input / 20% output came out to $5.40 / 1M tokens, vs $9 for Opus 4.8 and $10 for GPT-5.5 @jaminballArtificial Analysis estimated $0.94 average cost per Intelligence Index task, versus $1.04 for GPT-5.6 Sol and $1.80 for Opus 4.8 @ArtificialAnlysEarly live serving observations: ~28 tok/s via Moonshot API on OpenRouter @scaling01, and another observer saw 26 tok/s, calling it slower than Opus and speculating that speculative decoding wasn’t yet enabled @nrehiew_, @nrehiew_Moonshot’s blog reportedly recommends deployment on supernode configurations with 64+ accelerators for best inference efficiency @teortaxesTexvLLM said Moonshot contributed a KDA prefix caching implementation directly to vLLM, with support available day 0 for official release @vllm_projectMoonshot’s KDA contribution was cited as important because KDA breaks assumptions behind conventional prefix caching, so upstream runtime changes were required @vllm_projectBenchmarks and evalsMoonshot’s official benchmarking message, as summarized by others, positioned K3 behind only Claude Fable 5 and GPT-5.6 Sol among tested models, and ahead of Claude Opus 4.8 @scaling01, @Yuchenj_UWOne cited number: 1687 on GDPval-AA v2, above Opus 4.8 and behind GPT-5.6 Sol at 1747.8 in that comparison @scaling01Artificial Analysis’ independent numbers:AA Intelligence Index: 57GDPval v2 Elo: 1668AutomationBench-AA: 53%, #1AA-Briefcase Elo: 1547AA-Omniscience: +18, with accuracy 46% vs 33% on K2.6, but hallucination rate worsening to 51% from 39% @ArtificialAnlys, @ArtificialAnlysAA also reported 132M output tokens consumed for K3 across the Intelligence Index, versus 166M for K2.6, i.e. 21% reduction while gaining 13 index points @ArtificialAnlysArena’s frontend result was especially prominent because it is a pairwise human-preference arena, not just a static benchmark, and K3’s #1 frontend rank became one of the main launch headlines @arenaCommunity posts also highlighted strong results on kernel optimization tasks, with some saying K3 was matching or beating Fable in certain kernel/codegen settings @nrehiew_, @scaling01One benchmark caveat came from ProgramBench author Ofir Press, who said Kimi used a metric they do not recommend: averaging implementation percentage rather than counting fully working programs, which can overstate usefulness @OfirPress, @OfirPressFacts vs opinionsFacts / directly sourced claimsKimi K3 is officially announced by Moonshot @Kimi_MoonshotOfficially disclosed specs include 2.8T params, 1M context, native multimodal input, KDA, AttnRes, open weights by July 27 @Kimi_MoonshotArtificial Analysis independently scored K3 at 57 Intelligence Index, with detailed task, cost, token, and benchmark data @ArtificialAnlysArena independently ranked K3 #1 in Frontend Code Arena and later reported its 76% pairwise win rate @arena, @arenavLLM confirmed Moonshot contributed runtime support for KDA prefix caching @vllm_projectOpinions / interpretations“DeepSeek moment,” “beginning of the US-China AI race,” and “everything changed” are editorial interpretations from observers, not established facts @kimmonismus, @scaling01, @kimmonismusClaims that K3 “beats GPT-5.6 Sol o

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源
月之暗面发布新版Kimi模型引发担忧
TechCrunch AI(RSS)原文
Kimi团队再次发布最大中文模型,接近3T参数
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)原文

相似阅读

另一事件,读法相近