Cloudflare发布Clef决策模型:返回类型化概率而非文本
Cloudflare Releases Clef and Clef-flash: Open-Weight Decision Models That Return Typed Probabilities Instead of Text
决策模型范式值得关注,Clef直接输出带校准的概率分布,比传统LLM解析更稳定,适合Agent工具调用场景,建议关注其工程落地效果。
Cloudflare has released Clef and Clef-flash, the first models trained by its Workers AI team. They are decision models, not chatbots. Each reads an input state and a schema of typed questions. It returns a probability for every allowed answer, with no free-form text. Both are open-weight under Apache 2.0 and compatible with TypeSafe AI’s Jev API.
Cloudflare 发布了 Clef 和 Clef-flash,这是其 Workers AI 团队训练的首批模型。它们是决策模型,而非聊天机器人。每个模型读取输入状态和结构化问题模式,并为每个允许的答案返回概率,不生成自由文本。两者均采用 Apache 2.0 开源权重,并兼容 TypeSafe AI 的 Jev API。
Is it deployable? Yes, Both models run today on Workers AI, and the weights are on Hugging Face for self-hosting.
可以部署吗?可以。这两个模型目前均可在 Workers AI 上运行,且权重已托管在 Hugging Face 上以支持自托管。
What a Decision Model Does
决策模型的作用
An LLM generates tokens one at a time, and its output still needs parsing. A decision model only answers a fixed set of questions about an input. Clef supports 3 question types:
大语言模型(LLM)一次生成一个 token,其输出仍需解析。而决策模型仅针对输入回答一组固定的问题。Clef 支持三种问题类型:
- noul: yes/no, returns the probability of yes.
- choice: picks 1 named option, with per-option probabilities and a confidence value.
- score: rates against an ordered rubric, returning a probability-weighted score.
- noul:是/否,返回“是”的概率。
- choice:选择一个命名选项,并提供各选项的概率及置信度值。
- score:根据有序评分标准进行评级,返回概率加权的分数。
On Workers AI, 1 request carries up to 64 questions and up to 4 images.
在 Workers AI 上,单次请求最多可包含 64 个问题及最多 4 张图片。
TypeSafe AI launched Jev, its first ‘System One’ model, on September 15, 2026. Open alternatives like Kev-9B and Laya followed. Clef uses the same System One API. Switching from Jev means changing the endpoint and model name.
TypeSafe AI 于 2026 年 9 月 15 日推出了其首个‘System One’模型 Jev。Kev-9B 和 Laya 等开源替代方案随后跟进。Clef 使用相同的 System One API。从 Jev 切换至 Clef 只需更改端点和模型名称。
How Clef Works
Clef 的工作原理
Clef is post-trained from Qwen3.8-27B, and Clef-flash from Qwen3.5-9B. Both keep the backbone’s vision encoder.
Clef 基于 Qwen3.8-27B 进行后训练,Clef-flash 基于 Qwen3.5-9B 进行后训练。两者均保留了主干网络的视觉编码器。
Inference has 2 stages. The backbone first runs a single prefill-only pass over the state and questions. A small transformer, the joint schema head, then reads the final hidden states. It routes evidence to each question, lets fields cross-attend, and scores all options jointly. A per-question softmax turns logits into probabilities.
推理过程分为两个阶段。首先,主干网络对状态和问题执行单次仅预填充(prefill-only)的前向传播。随后,一个小型 Transformer——联合模式头(joint schema head)——读取最终的隐藏状态。它将证据路由至每个问题,使字段相互交叉注意,并联合对所有选项进行评分。每道题目的 softmax 函数将 logits 转换为概率。
Training froze both backbones and jointly optimized the routing head with rank-256 low-rank adapters. The loss pairs label-smoothed cross-entropy with a Brier loss for calibration. A secondary objective, Reinforcement Learning for Calibrated Decisions (RLCD), gives partial credit to adjacent ordinal choices.
训练过程中冻结了两个主干网络,并使用秩为 256 的低秩适配器(low-rank adapters)联合优化路由头。损失函数结合了标签平滑交叉熵与用于校准的 Brier 损失。次要目标采用校准决策强化学习(RLCD),对相邻的顺序选择给予部分奖励。
Interactive Explainer
交互式解释器
Benchmarks: Where Clef Wins and Where It Does Not
基准测试:Clef 的优势与不足
On Cloudflare’s 10-benchmark shortlist from the Decision Index 0.2.1 suite, a Clef model scored highest on 7.
在 Decision Index 0.2.1 套件中 Cloudflare 选出的 10 个基准测试短名单中,Clef 模型在其中 7 个测试中得分最高。
- BANKING77 (macro-F1): Clef 94.20 vs Jev 79.74.
- CLINC150+OOS (macro-F1): Clef 97.43 vs Jev 89.27.
- Home appliances (case exact): Clef-flash 97.73 vs Jev 52.27.
- BANKING77(宏 F1 分数):Clef 94.20 vs Jev 79.74。
- CLINC150+OOS(宏 F1 分数):Clef 97.43 vs Jev 89.27。
- 家用电器(精确匹配案例):Clef-flash 97.73 vs Jev 52.27。
Jev keeps clear leads elsewhere. The full model card shows Jev ahead on GPQA Diamond (78.3 vs 48.0). It also leads MMLU-Pro (82.7 vs 65.9) and BBH (92.9 vs 73.7).
Jev 在其他领域仍保持明显领先。完整的模型卡片显示,Jev 在 GPQA Diamond(78.3 vs 48.0)、MMLU-Pro(82.7 vs 65.9)和 BBH(92.9 vs 73.7)上均领先。
On TypeSafe’s own workflow evals, Clef beat Jev in 3 of 4 areas, by small margins. Invoice processing was 64.7 vs 61.8, customer service 76.3 vs 76.0, and security incidents 62.9 vs 61.7. Jev leads agent trace observability, 71.6 vs 68.5.
在 TypeSafe 自身的工作流评估中,Clef 在 4 个领域中的 3 个以微弱优势击败 Jev。发票处理为 64.7 对 61.8,客户服务为 76.3 对 76.0,安全事件为 62.9 对 61.7。Jev 在代理追踪可观测性方面领先,得分为 71.6 对 68.5。
In Cloudflare’s threat intelligence workflow, Clef classified a domain in 2.2 seconds. gpt-oss-120b took 4.7 seconds.
在 Cloudflare 的威胁情报工作流中,Clef 用 2.2 秒对域名进行分类。gpt-oss-120b 耗时 4.7 秒。
All numbers are vendor-reported, with no independent replication yet.
所有数据均由供应商报告,目前尚无独立复现结果。
Feature Comparison
功能对比
| Feature | Clef | Clef-flash | Jev | Kev-9B | Laya |
|---|---|---|---|---|---|
| Developer | Cloudflare | Cloudflare | TypeSafe AI | Jared Palmer | Convai Innovations |
| Size | 27B | 9B | Not disclosed | 9B + 45.4M LoRA | 421M |
| Backbone | Qwen3.8-27B | Qwen3.5-9B | Not disclosed | Qwen3.5-9B-Base | ModernBERT-large |
| Weights | Apache 2.0 | Apache 2.0 | Hosted API | Apache 2.0 | Apache 2.0 |
| Image input | Yes | Yes | No | No | No |
| Context | 65,536 | 65,536 | 32K (per Cloudflare) | 65,536 (8,192 validated) | 512 (English) |
| Median latency* | 209.3 ms | 38.8 ms | 524.1 ms | 51.4 ms | 5.8 ms |
| Hosted price (input) | $0.24/M | $0.09/M | $0.042/M | Self-host | Self-host |
| 功能 | Clef | Clef-flash | Jev | Kev-9B | Laya |
|---|---|---|---|---|---|
| 开发者 | Cloudflare | Cloudflare | TypeSafe AI | Jared Palmer | Convai Innovations |
| 规模 | 27B | 9B | 未披露 | 9B + 45.4M LoRA | 421M |
| 骨干网络 | Qwen3.8-27B | Qwen3.5-9B | 未披露 | Qwen3.5-9B-Base | ModernBERT-large |
| 权重许可 | Apache 2.0 | Apache 2.0 | 托管 API | Apache 2.0 | Apache 2.0 |
| 图像输入 | 支持 | 支持 | 不支持 | 不支持 | 不支持 |
| 上下文窗口 | 65,536 | 65,536 | 32K(据 Cloudflare) | 65,536(8,192 已验证) | 512(英文) |
| 中位延迟* | 209.3 ms | 38.8 ms | 524.1 ms | 51.4 ms | 5.8 ms |
| 托管价格(输入) | $0.24/M | $0.09/M | $0.042/M | 自托管 | 自托管 |
*Cloudflare’s internal Decision Index run. All 5 implement the System One API. Sources: Clef docs, Clef-flash docs, Clef card, Jev post, Kev-9B card, Laya card.
*Cloudflare 内部决策指数运行结果。所有 5 个模型均实现 System One API。来源:Clef 文档、Clef-flash 文档、Clef 卡片、Jev 帖子、Kev-9B 卡片、Laya 卡片。
Deployment and Fine-Tuning
部署与微调
Both models are callable through the Workers AI binding (env.AI.run ), the REST API, or AI Gateway. For self-hosting, the model cards list testing on a single H200 with BF16 weights.
这两个模型均可通过 Workers AI 绑定(env.AI.run)、REST API 或 AI Gateway 调用。对于自托管场景,模型卡片列出了在单张 H200 GPU 上使用 BF16 权重的测试结果。
Cloudflare also announced a reinforcement learning service for tuning Clef on private data. It starts with Cloudflare’s forward-deployed engineers, with a self-serve platform later. The pipeline combines AI Gateway, Workers AI, Containers and a new Trainer component. Teams can apply via the design partner form.
Cloudflare 还宣布了一项强化学习服务,用于在私有数据上微调 Clef。该服务最初由 Cloudflare 的前端部署工程师提供,随后将推出自助服务平台。该流水线结合了 AI Gateway、Workers AI、Containers 以及一个新的 Trainer 组件。团队可通过设计合作伙伴申请表进行申请。
Key Takeaways
关键要点
- Clef (27B) and Clef-flash (9B) are Apache 2.0, Jev-compatible decision models.
- Median latency: 209.3 ms for Clef, 38.8 ms for Clef-flash, 524.1 ms for Jev.
- Clef reads text, JSON, images and video within a 64K-token context window.
- Jev still leads on knowledge-heavy tests like GPQA Diamond, MMLU-Pro and BBH.
- An RL fine-tuning service starts with Cloudflare’s forward-deployed engineers.
- Clef (27B) 和 Clef-flash (9B) 是 Apache 2.0 许可的、与 Jev 兼容的决策模型。
- 中位延迟:Clef 为 209.3 ms,Clef-flash 为 38.8 ms,Jev 为 524.1 ms。
- Clef 可在 64K token 的上下文窗口内读取文本、JSON、图像和视频。
- Jev 仍在 GPQA Diamond、MMLU-Pro 和 BBH 等知识密集型测试中保持领先。
- 一项强化学习微调服务最初由 Cloudflare 的前端部署工程师启动。
Check out the Model weight, Demo and Technical details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
查看模型权重、演示和技术细节。所有功劳归于本项目的研究人员。此外,欢迎在 Twitter 上关注我们,并别忘了加入我们有 15 万+成员的 ML SubReddit 以及订阅我们的新闻通讯。等等!你在 Telegram 上吗?现在你也可以加入我们的 Telegram 群组。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力