跳到主内容
精选90Rohan Paul模型发布/更新多源精选 ×11

月之暗面发布Kimi K3:2.8万亿参数MoE,百万token上下文

Beijing-based Moonshot AI's Kimi K3 just dropped.

原文
推荐理由

国产模型新旗舰,2.8T参数MoE架构+百万上下文,技术报告亮点多,做长上下文或MoE的同学值得细读。

Beijing-based Moonshot AI's Kimi K3 just dropped.

  • 1 Mn token context window, natively multimodal. (i.e. 750,000 words of code or documentation in a single prompt.) - Total 2.8 trillion params, but activates only 16 of its 896 experts at a time. That is roughly 1.8% of the expert pool per token. - Some of the benchmarks put it Opus 4.8/ GPT 5.6 Sol / Fable 5 territory. - Its Delta Attention enables up to 6.3x faster decoding in million-token contexts - available via API at $3/$15 per million input/output tokens
  • K3’s 2.8T MXFP4 weights require roughly 1.4 TB before quantization metadata and runtime overhead.

At an assumed 115 GB usable per GB10 node (NVIDIA’s compact Grace Blackwell AI chip), 14–16 nodes is expected to hold the raw weights realistically. Including activations, KV cache and runtime memory etc. Moonshot itself recommends supernodes containing 64 or more accelerators.

Some very cool findings from technical report

  • K3 autonomously designed, optimized, and verified a working AI chip in a single 48-hour run—specifically to serve a smaller model built on K3’s own architecture.

The simulated chip reportedly reached 8,700+ tokens/second, contained 1.46 million standard cells, and fit within 4 mm².

  • An early version of K3 handled the majority of the kernel-optimization work used to develop K3 itself.
  • K3 built a GPU compiler from scratch. created MiniTriton, optimization passes, PTX generation, and runtime. Matched or beat Triton on some workloads and successfully trained nanoGPT end to end.
  • In a 15-hour autonomous run, Kimi K3 redesigned a production-scale training kernel and cut forward-plus-backward time from 283.6 ms to 114.4 ms.
  • For a 42-year semiconductor-industry report, Kimi says K3 performed 2,800+ web searches/fetches, 1,100+ terminal data pulls, processed 11,000+ pages, and recursively improved the work over 120+ rounds.
  • K3 reproduced a computational-astrophysics research workflow in roughly two hours, versus an estimated one to two weeks for an experienced researcher. It reviewed 20+ papers, evaluated 300+ equations of state, found inconsistencies in published formulas, and wrote 3,000+ lines of Python.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源
月之暗面发布新版Kimi模型引发担忧
TechCrunch AI(RSS)原文
Kimi团队再次发布最大中文模型,接近3T参数
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)原文

相似阅读

另一事件,读法相近