Aleph Alpha发布Kolibri:781亿参数英德双语MoE模型
Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model With Only 3.46B Active Parameters
新发布的MoE模型兼顾了极致效率与合规性,单卡即可部署百万上下文,对注重数据主权和私有化部署的团队极具参考价值。
Aleph Alpha has released Kolibri, an open-weight Mixture-of-Experts (MoE) language model built for German and English. Kolibri has 78.1B total parameters but activates only 3.46B, or 4.4%, per token. It accepts up to 1,048,576 tokens of context, lets users set reasoning effort per request, and ships under the Apache 2.0 license on Hugging Face. The target is sovereign deployment in regulated sectors such as public administration, industry and aerospace.
Aleph Alpha 发布了 Kolibri,这是一个专为德语和英语设计的开源权重混合专家(MoE)语言模型。Kolibri 拥有 781 亿总参数,但每个 token 仅激活 34.6 亿,即 4.4%。它支持高达 1,048,576 个 token 的上下文,允许用户为每次请求设置推理力度,并在 Hugging Face 上以 Apache 2.0 许可证发布。其目标是在公共行政、工业和航空航天等受监管领域实现主权部署。
Is it deployable? Yes. The FP8 checkpoint is about 78GB and runs on a single B200, B300 or H200, or on 2 H100 SXM5 GPUs, served through vLLM with dedicated Kolibri reasoning and tool-call parsers.
可以部署吗?可以。FP8 检查点约为 78GB,可在单个 B200、B300 或 H200 上运行,或在 2 块 H100 SXM5 GPU 上运行,通过 vLLM 提供服务,并配备专用的 Kolibri 推理和工具调用解析器。
What is Kolibri?
什么是 Kolibri?
Kolibri (Kolibri-1) is a bilingual English-German MoE transformer developed end to end by teams in Germany. According to the technical research report, Aleph Alpha team controlled the full pipeline: data, architecture, training infrastructure, post-training and evaluation. Training ran on infrastructure in Germany and Finland. The design targets the EU General-Purpose AI Code of Practice, the EU AI Act and GDPR. Aleph Alpha is a signatory of that Code, and its data pipeline redacts personal data before training.
Kolibri (Kolibri-1) 是一个双语英德 MoE Transformer,由德国的团队从头开发。根据技术研究报告,Aleph Alpha 团队控制了完整的流水线:数据、架构、训练基础设施、后训练和评估。训练在德国和芬兰的基础设施上进行。设计旨在符合欧盟通用人工智能行为准则、欧盟《人工智能法案》和 GDPR。Aleph Alpha 是该行为准则的签署方,其数据管道在训练前会对个人数据进行脱敏处理。
Architecture: Sparse Experts and Hybrid Attention
架构:稀疏专家与混合注意力
Kolibri stacks 50 transformer blocks with a model width of 2,560. Every MoE layer scores all 384 routed experts with a sigmoid router, sends each token to the top 6, and always runs 1 shared expert. Expert load is balanced with Exact Quantile Balancing and Load-Error Injection.
Kolibri 堆叠了 50 个 Transformer 块,模型宽度为 2,560。每个 MoE 层使用 sigmoid 路由器对所有 384 个路由专家进行评分,将每个 token 发送到前 6 个专家,并始终运行 1 个共享专家。专家负载通过精确分位数平衡和负载误差注入进行平衡。
Attention uses grouped-query attention with 48 query heads and 4 KV heads. Every fifth block uses full attention without positional encoding. The other 40 blocks use sliding-window attention over the 512 preceding tokens, with RoPE. Sliding-window layers hold a fixed-size KV cache, so only 10 layers grow with context length. At matched compute, Aleph Alpha team reports the hybrid supports sequences 4 times longer than a full-attention model.
注意力机制采用分组查询注意力,包含 48 个查询头和 4 个 KV 头。每第五个块使用不带位置编码的全注意力。其余 40 个块使用滑动窗口注意力,覆盖前 512 个 token,并采用 RoPE。滑动窗口层持有固定大小的 KV 缓存,因此只有 10 个层随上下文长度增长。在计算量匹配的情况下,Aleph Alpha 团队报告称,这种混合模式支持的序列长度是全注意力模型的 4 倍。
A Tokenizer Built for German
专为德语构建的分词器
The 128,000-token vocabulary is trained with UniBPE, which builds merges like BPE but scores each merge by Unigram loss. On German text it reaches 4.90 bytes per token, versus 4.35 for the GPT-5 tokenizer. That means 11.2% fewer tokens on German web text. In English, Kolibri reaches 4.58 bytes per token against 4.67 for GPT-5.
128,000 的词表是使用 UniBPE 训练的,它像 BPE 一样构建合并操作,但通过 Unigram 损失对每次合并进行评分。在德语文本上,它达到每个 token 4.90 字节,而 GPT-5 分词器为 4.35 字节。这意味着在德语网络文本上减少了 11.2% 的 token 数量。在英语方面,Kolibri 达到每个 token 4.58 字节,而 GPT-5 为 4.67 字节。
Training: 24T Tokens, Then SFT and RL
训练:24T Token,然后是 SFT 和 RL
Pre-training covered 20T tokens on 768 NVIDIA B200 GPUs, followed by 3.44T mid-training tokens at 65,536 sequence length. A 201B-token long-context stage then trained on 262,144-token sequences. Aleph Alpha added more than 2T German tokens it curated from the web or generated synthetically. Post-training combined supervised fine-tuning, mixed with MergeMix, with reinforcement learning on more than 1.2M internal tasks. The Merlin-Arthur protocol trains the model to abstain when retrieved context does not support an answer.
预训练在 768 块 NVIDIA B200 GPU 上覆盖了 20T token,随后在序列长度 65,536 下进行了 3.44T mid-training token 的训练。接着是一个 201B-token 的长上下文阶段,在 262,144-token 序列上进行训练。Aleph Alpha 添加了超过 2T 个从网络中整理或合成生成的德语 token。后训练结合了监督微调(与 MergeMix 混合)以及在超过 120 万个内部任务上的强化学习。Merlin-Arthur 协议训练模型在检索到的上下文不支持答案时选择拒答。
Interactive Explainer: How Kolibri Works
交互式解释器:Kolibri 的工作原理
Benchmarks
基准测试
Aleph Alpha team evaluated every model with the same eval-framework setup. In English, Kolibri leads GPQA Diamond (84.3), AIME 2025 (96.9) and AIME 2026 (96.0). It ties Qwen3.5 35B-A3B on the English agentic average at 63.4. It trails on BFCL v4, scoring 61.4 against 70.5 for Qwen3.5. The dense Qwen3.8 27B scores higher overall (80.2 EN, 79.9 DE), but activates about 8 times more parameters per token. Against its internal predecessor Kolibri Origin, Kolibri decodes about 2.7 times more text per GPU while scoring 21.4 points higher in English.
Aleph Alpha 团队使用相同的评估框架设置对每个模型进行了评估。在英语方面,Kolibri 在 GPQA Diamond(84.3)、AIME 2025(96.9)和 AIME 2026(96.0)上领先。它在英语智能体平均值上与 Qwen3.5 35B-A3B 持平,均为 63.4。在 BFCL v4 上它落后于对手,得分为 61.4,而 Qwen3.5 为 70.5。稠密模型 Qwen3.8 27B 的整体得分更高(英语 80.2,德语 79.9),但每个 token 激活的参数数量约为前者的 8 倍。与其内部前身 Kolibri Origin 相比,Kolibri 每块 GPU 解码的文本量约为前者的 2.7 倍,且英语得分高出 21.4 分。
Kolibri vs Closest Competitors
Kolibri 与最接近的竞争者对比
| Feature | Kolibri-1 | Qwen3.6 35B-A3B | Nemotron 3 Super | Mistral Small 4 |
|---|---|---|---|---|
| Developer | Aleph Alpha (Germany) | Alibaba Qwen | NVIDIA | Mistral AI (France) |
| Total / active params | 78.1B / 3.46B | ~35B / 3B | 120B / 12B | 119B / 6.5B |
| Architecture | MoE, sliding-window + full attention | MoE, Gated DeltaNet hybrid | Mamba-2 + MoE + attention | MoE |
| Max context | 1,048,576 tokens | 262,144 native, ~1M with YaRN | 1M tokens | 256k tokens |
| Reasoning control | none / low / medium / high | Thinking on / off | On / off, low-effort mode | none / high |
| License | Apache 2.0 | Apache 2.0 | NVIDIA Nemotron Open Model License | Apache 2.0 |
| Overall score (EN / DE) | 75.5 / 70.8 | 71.4 / 67.3 | 73.0 / 67.9 | 63.1 / 61.4 |
| Agentic avg (EN) | 63.4 | 62.1 | 54.9 | 40.7 |
| Industry RAG avg (DE) | 67.5 | 65.8 | 57.6 | 53.4 |
| Code avg (EN) | 89.3 | 87.7 | 88.3 | 82.0 |
| 特性 | Kolibri-1 | Qwen3.6 35B-A3B | Nemotron 3 Super | Mistral Small 4 |
|---|---|---|---|---|
| 开发者 | Aleph Alpha (德国) | Alibaba Qwen | NVIDIA | Mistral AI (法国) |
| 总参数 / 活跃参数 | 78.1B / 3.46B | ~35B / 3B | 120B / 12B | 119B / 6.5B |
| 架构 | MoE,滑动窗口 + 全注意力 | MoE,Gated DeltaNet 混合 | Mamba-2 + MoE + 注意力 | MoE |
| 最大上下文 | 1,048,576 tokens | 原生 262,144,使用 YaRN 约 1M | 1M tokens | 256k tokens |
| 推理控制 | 无 / 低 / 中 / 高 | Thinking on / off | On / off,低努力模式 | 无 / 高 |
| 许可证 | Apache 2.0 | Apache 2.0 | NVIDIA Nemotron Open Model License | Apache 2.0 |
| 综合得分 (EN / DE) | 75.5 / 70.8 | 71.4 / 67.3 | 73.0 / 67.9 | 63.1 / 61.4 |
| 智能体平均值 (EN) | 63.4 | 62.1 | 54.9 | 40.7 |
| 行业 RAG 平均值 (DE) | 67.5 | 65.8 | 57.6 | 53.4 |
| 代码平均值 (EN) | 89.3 | 87.7 | 88.3 | 82.0 |
Specs: model cards for Kolibri-1, Qwen3.6-35B-A3B, Nemotron 3 Super and Mistral Small 4.
规格:Kolibri-1、Qwen3.6-35B-A3B、Nemotron 3 Super 和 Mistral Small 4 的模型卡片。
How to Deploy Kolibri
如何部署 Kolibri
Install the aleph-alpha-inference package, then serve with vLLM:
安装 aleph-alpha-inference 包,然后使用 vLLM 提供服务:
pip install 'aleph-alpha-inference>=1'
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
--reasoning-parser kolibri1 --tool-call-parser kolibri1 \
--enable-auto-tool-choiceThe default context is 262,144 tokens. Pass --max-model-len 1048576 with a max_position_embeddings override for the full 1M window. Set reasoning_effort through chat_template_kwargs. Aleph Alpha recommends temperature 1.0, top-p 0.97 and top-k 128.
默认上下文长度为 262,144 tokens。传递 --max-model-len 1048576 并覆盖 max_position_embeddings 以启用完整的 1M 窗口。通过 chat_template_kwargs 设置 reasoning_effort。Aleph Alpha 推荐 temperature 为 1.0,top-p 为 0.97,top-k 为 128。
Key Takeaways
关键要点
- 78.1B total, 3.46B active: each token uses 6 of 384 routed experts plus 1 shared expert.
- 40 sliding-window and 10 full-attention layers keep a 1M-token context affordable.
- Trained on 24T tokens, with German above 20% of the mix.
- Top Overall score in English (75.5) and German (70.8) among the 12 MoE models Aleph Alpha compared.
- Apache 2.0 FP8 weights, serveable on a single B200 or H200 via vLLM.
- 总计78.1B,活跃参数3.46B:每个token使用384个路由专家中的6个以及1个共享专家。
- 40个滑动窗口层和10个全注意力层使得支持1M token上下文变得经济可行。
- 在24T token上训练,其中德语占比超过20%。
- 在Aleph Alpha对比的12个MoE模型中,英语(75.5)和德语(70.8)的总体得分最高。
- Apache 2.0协议的FP8权重,可通过vLLM在单张B200或H200上部署服务。
Check out the Model Weights and Technical Report. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
查看模型权重和技术报告。所有功劳归于该项目的研究人员。此外,欢迎在Twitter上关注我们,别忘了加入我们拥有15万+成员的ML SubReddit,并订阅我们的新闻通讯。等等!你在Telegram上吗?现在你也可以加入我们的Telegram群组。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力