Cloudflare发布开源决策模型Clef及RL微调平台
Introducing Clef: our open-source decision models, and new RL fine-tuning platform
决策模型是Agent架构的关键组件,Clef提供了开源替代方案并附带RL微调能力,适合需要低成本、高确定性分类的开发者参考。
Over the last few weeks, there has been lots of buzz around decision models such as Typesafe AI’s Jev System One model. While classifier models have been around for some time, Jev introduces a new decision model concept into the world of AI — a model that produces bounded structured outputs cheaply, quickly and consistently that can be added into a workflow when a decision is required. These models are capable enough to work over any set of inputs without constantly retraining the model to incorporate new classification categories. This contrasts with the world of Large Language Models (LLMs), which are largely non-deterministic, but are open-ended enough to reason and generate text and tool calls for agentic workloads.
在过去几周里,围绕决策模型(如 Typesafe AI 的 Jev System One 模型)引发了大量热议。虽然分类模型已经存在了一段时间,但 Jev 将一种新的决策模型概念引入了 AI 领域——这是一种能够以低成本、快速且一致的方式生成有界结构化输出的模型,可在需要做出决策时集成到工作流中。这些模型具备足够的处理能力,可以在任何输入集上运行,而无需为了纳入新的分类类别而不断重新训练模型。这与大型语言模型(LLMs)的世界形成对比,后者在很大程度上是非确定性的,但具有足够的开放性,能够为智能体工作负载进行推理并生成文本和工具调用。
Today, we’re releasing two Cloudflare-trained decision models, Clef and Clef-flash, hosted on Workers AI. Clef is currently the leader when evaluated against the Jev Decision Index, you can view full results on the live benchmark demo site. These models are smarter, faster, and fully Jev-API compatible, so you can experiment with these hosted models easily. We’re fully open-sourcing these models on Hugging Face under an Apache 2.0 license for you to run locally and experiment with yourselves.
今天,我们发布了两个由 Cloudflare 训练的决策模型 Clef 和 Clef-flash,它们托管在 Workers AI 上。在对 Jev Decision Index 进行评估时,Clef 目前处于领先地位,您可以在实时基准测试演示网站上查看完整结果。这些模型更智能、更快,并且完全兼容 Jev-API,因此您可以轻松尝试这些托管模型。我们已将这些模型在 Hugging Face 上完全开源,采用 Apache 2.0 许可证,供您本地运行并自行实验。
Lastly, we’re excited to debut our new reinforcement learning (RL) product, which allows customers to fine-tune Clef to suit their use cases as well.
最后,我们兴奋地推出全新的强化学习(RL)产品,允许客户对 Clef 进行微调以适应其用例。
What is a decision model?
什么是决策模型?
A decision model makes classifications to help agents decide how to act, based on certain probabilities. For example, you can pass in a customer support message (inputs) and ask if it is urgent and which team should handle it. A decision model will return typed answers with probabilities (outputs), which your code can use to route the ticket, trigger an escalation, or defer to a human. This means that a human does not necessarily need to be in the loop for agentic decisions anymore — agents can programmatically gather context, make decisions, and take actions on tasks, or defer to a human when needed.
决策模型基于特定概率进行分类,以帮助智能体决定如何行动。例如,您可以传入一条客户支持消息(输入),并询问它是否紧急以及应由哪个团队处理。决策模型将返回带有概率的类型化答案(输出),您的代码可以利用这些答案来路由工单、触发升级流程或在必要时转交人工处理。这意味着在智能体决策中,人类不一定需要始终参与循环——智能体可以程序化地收集上下文、做出决策并对任务采取行动,或在需要时转交人工处理。
Specifically at Cloudflare, we’ve been testing our new Clef model on our Threat Intelligence team to help us classify website domains. By giving a domain to Clef (with Browser Run) it can quickly identify categories that the domain falls under — for example, it might classify a domain with a 95% chance it is a fashion website, 85% ecommerce, <1% phishing, etc. This classification took our Clef model 2.2s to fetch, render, and classify the website. In contrast, our fastest general LLM gpt-oss-120b took 4.7s in the same workflow, and only returned two classifications. As a user, you can imagine how a 2x savings in latency and results can help us improve our threat intelligence workflows and be faster in identifying malicious or legitimate domains. Generalize this to any use case where you need to make quick programmatic decisions, and you unlock powerful agentic workflows that are able to autonomously decide, reason, and execute.
具体而言,在 Cloudflare,我们一直在威胁情报团队中测试我们的新 Clef 模型,以帮助我们对网站域名进行分类。通过将域名提供给 Clef(配合 Browser Run),它可以快速识别该域名所属的类别——例如,它可能以 95% 的概率将其分类为时尚网站,85% 的概率为电子商务网站,<1% 的概率为钓鱼网站等。我们的 Clef 模型完成获取、渲染和分类网站的过程耗时 2.2 秒。相比之下,在我们最快的通用大语言模型 gpt-oss-120b 执行相同工作流时耗时 4.7 秒,且仅返回两个分类结果。作为用户,你可以想象延迟和结果的节省如何帮助我们改进威胁情报工作流,并更快地识别恶意或合法域名。将此推广到任何需要快速进行程序化决策的用例,你将解锁强大的智能体工作流,使其能够自主做出决策、推理和执行。
In music theory, a clef is a symbol placed at the beginning of a musical staff that assigns specific pitch names to the lines and spaces. A decision model is analogous to a music clef because it helps define the domain of the context and the subsequent notes (actions) that follow it. We chose Clef as the name of our family of decision models, as it serves similar purposes, and the CF hearkens to Cloudflare.
在音乐理论中,谱号是放置在乐谱开头的一个符号,用于将特定的音高名称分配给五线谱上的线和间。决策模型类似于音乐谱号,因为它有助于定义上下文的领域以及随后出现的音符(动作)。我们将我们的决策模型系列命名为 Clef,因为它具有类似的目的,且“CF”呼应了 Cloudflare。
How is Clef different from other decision models?
Clef 与其他决策模型有何不同?
Although the market is getting increasingly saturated with decision models, Clef has some unique properties that make us excited to release it to the public. First, it has a vision encoder so it’s able to take in images and classify visual content. This is different from Jev, which only does text classification today. Secondly, our model has a 64k context window (compared to Jev’s 32k), which allows users to squeeze more input state for the model to classify against.
尽管市场上充斥着越来越多的决策模型,但 Clef 具有一些独特的特性,使我们兴奋地向公众发布它。首先,它配备了一个视觉编码器,因此能够接收图像并对视觉内容进行分类。这与 Jev 不同,后者目前仅支持文本分类。其次,我们的模型拥有 64k 的上下文窗口(相比之下 Jev 为 32k),这允许用户向模型输入更多的状态以供其进行分类判断。
Third, our model is accurate and powerful, scoring competitively against other decision models on the market across various quality benchmarks. We shortlisted some evaluations below that are important for decision-making as defined by the Jev Decision Index and scored some of the more popular models on the market for it. Check out the table below for benchmarks, or view the scores on our live decision index demo site:
第三,我们的模型准确且强大,在各种质量基准测试中与市场上的其他决策模型相比具有竞争力。根据 Jev Decision Index 定义的决策相关标准,我们筛选了一些评估指标,并对市场上一些更流行的模型进行了评分。请查看下方的基准测试表,或在我们的实时决策指数演示网站上查看分数:
| Benchmark | Clef | Clef-flash | Jev | DiffusionGemma Jev | Kev 9B | Laya |
|---|---|---|---|---|---|---|
| BFCL · case exact | 98.47 | 98.76 | 95.75 | 96.52 | 94.51 | 38.13 |
| ToolRet · nDCG@10 | 69.19 | 66.43 | 65.28 | 61.21 | 64.26 | 12.69 |
| API-Bank · accuracy | 91.93 | 93.11 | 88.19 | 83.66 | 56.30 | 11.41 |
| Home appliances · case exact | 82.95 | 97.73 | 52.27 | 42.05 | 25.00 | 0.00 |
| When2Call · accuracy | 72.37 | 65.58 | 80.97 | 75.44 | 49.62 | 11.94 |
| BANKING77 · macro-F1 | 94.20 | 90.93 | 79.74 | 74.28 | 84.83 | 14.29 |
| CLINC150+OOS · macro-F1 | 97.43 | 66.77 | 89.27 | 83.49 | 79.03 | 3.19 |
| BRIGHT · nDCG@10 | 45.91 | 39.26 | 47.52 | 42.94 | 38.53 | 19.90 |
| Amazon ESCI · macro-F1 | 57.48 | 57.39 | 55.21 | 53.37 | 49.22 | 24.40 |
| PhishNChips · accuracy | 79.60 | 75.05 | 62.55 | 85.35 | 50.75 | 50.15 |
| 基准测试 | Clef | Clef-flash | Jev | DiffusionGemma Jev | Kev 9B | Laya |
|---|---|---|---|---|---|---|
| BFCL · case exact | 98.47 | 98.76 | 95.75 | 96.52 | 94.51 | 38.13 |
| ToolRet · nDCG@10 | 69.19 | 66.43 | 65.28 | 61.21 | 64.26 | 12.69 |
| API-Bank · accuracy | 91.93 | 93.11 | 88.19 | 83.66 | 56.30 | 11.41 |
| Home appliances · case exact | 82.95 | 97.73 | 52.27 | 42.05 | 25.00 | 0.00 |
| When2Call · accuracy | 72.37 | 65.58 | 80.97 | 75.44 | 49.62 | 11.94 |
| BANKING77 · macro-F1 | 94.20 | 90.93 | 79.74 | 74.28 | 84.83 | 14.29 |
| CLINC150+OOS · macro-F1 | 97.43 | 66.77 | 89.27 | 83.49 | 79.03 | 3.19 |
| BRIGHT · nDCG@10 | 45.91 | 39.26 | 47.52 | 42.94 | 38.53 | 19.90 |
| Amazon ESCI · macro-F1 | 57.48 | 57.39 | 55.21 | 53.37 | 49.22 | 24.40 |
| PhishNChips · accuracy | 79.60 | 75.05 | 62.55 | 85.35 | 50.75 | 50.15 |
We also ran benchmarks across Typesafe’s own eval suite and our Clef models fared well, beating Jev in 3 out of 4 areas. Notably, our Clef-flash performs exceptionally well, given how much faster it is.
我们还在 Typesafe 自己的评估套件中运行了基准测试,我们的 Clef 模型表现良好,在 4 个领域中的 3 个领域超越了 Jev。值得注意的是,考虑到 Clef-flash 的速度要快得多,其表现尤为出色。
| Workflow | Clef | Clef-flash | Jev |
|---|---|---|---|
| Invoice processing | 64.7 | 57.1 | 61.8 |
| Customer service | 76.3 | 77 | 76.0 |
| Security incidents | 62.9 | 61.7 | 61.7 |
| Agent trace observability | 68.5 | 69.8 | 71.6 |
| 工作流 | Clef | Clef-flash | Jev |
|---|---|---|---|
| 发票处理 | 64.7 | 57.1 | 61.8 |
| 客户服务 | 76.3 | 77 | 76.0 |
| 安全事件 | 62.9 | 61.7 | 61.7 |
| 代理追踪可观测性 | 68.5 | 69.8 | 71.6 |
Across the 43 eval benchmarks that we ran, our Clef models beat the decision models on latency (except for Laya which is very fast but trades off quality in the benchmarks above):
在我们运行的 43 个评估基准中,我们的 Clef 模型在延迟方面优于决策模型(Laya 除外,它非常快,但在上述基准测试中牺牲了质量):
| Benchmark | Clef | Clef-flash | Jev | DiffusionGemma Jev | Kev-9B | Laya |
|---|---|---|---|---|---|---|
| Median latency · ms | 209.3 | 38.8 | 524.1 | 84.4 | 51.4 | 5.8 |
| p95 latency · ms | 238.6 | 122.4 | 536.0 | 211.2 | 187.9 | 222.5 |
| 基准测试 | Clef | Clef-flash | Jev | DiffusionGemma Jev | Kev-9B | Laya |
|---|---|---|---|---|---|---|
| 中位延迟 · ms | 209.3 | 38.8 | 524.1 | 84.4 | 51.4 | 5.8 |
| p95 延迟 · ms | 238.6 | 122.4 | 536.0 | 211.2 | 187.9 | 222.5 |
On top of the latency benefits from the model itself, our Clef models are hosted on Workers AI. Because they are hosted on Cloudflare’s infrastructure, we’re able to take advantage of our GPUs at the edge, leading to low network latency and faster decisions. This means that you could put Clef into the hot path for agents to make decisions and combine that with one of our LLMs on Workers AI to take action.
除了模型本身带来的延迟优势外,我们的 Clef 模型还托管在 Workers AI 上。由于它们托管在 Cloudflare 的基础设施上,我们能够利用边缘的 GPU,从而实现低网络延迟和更快的决策。这意味着你可以将 Clef 放入代理决策的热路径中,并将其与 Workers AI 上的其中一个 LLM 结合以采取行动。
Clef also produces strictly typed outputs similar to Jev and is fully API-compatible, so you can make the swap extremely easily. The larger Clef model is your more powerful precision model, while the Clef-Flash model is great for latency-critical decisions. The models are enterprise-ready with our guarantee that we don’t read, store, or train on your requests or responses (unless you want to use our fine-tuning product, which we go into below). You can get started with the Clef models today, starting with our developer documentation or play around with the open-source model on the Hugging Face repo.
Clef 还会生成与 Jev 类似的严格类型输出,并且完全兼容 API,因此您可以极其轻松地进行替换。体积更大的 Clef 模型是功能更强大的高精度模型,而 Clef-Flash 模型则非常适合对延迟敏感的场景。这些模型已具备企业级就绪能力,我们承诺不会读取、存储或基于您的请求或响应进行训练(除非您希望使用我们的微调产品,下文将对此进行介绍)。您可以从今天开始使用 Clef 模型,从我们的开发者文档入手,或在 Hugging Face 仓库中试用开源模型。
If you’d like help tuning Clef for a specific workload, we are also offering fine-tuning services — first as a hands-on partner with our forward-deployed engineer (FDE) team, and then later as a self-serve fine-tuning platform for customers to train and redeploy the model onto Cloudflare.
如果您需要帮助针对特定工作负载调整 Clef,我们还提供微调服务——首先以现场部署工程师(FDE)团队的身份作为合作伙伴提供指导,随后将推出自助式微调平台,供客户在 Cloudflare 上训练并重新部署模型。
How we trained Clef
我们如何训练 Clef
In the same week that Jev came out, we posted about some experiments we had with our own homegrown decision model. Our demo goes into how we adapted the DiffusionGemma model to output deterministic probabilities by exposing the logprobs that are generated by a large language model. Our initial approach built upon independent research by Matt Mastracci, who has been active in the machine learning (ML) community with sharing new ideas and pull requests to vLLM inference engine to make DiffusionGemma support stronger.
在 Jev 发布的那一周,我们分享了关于自有决策模型的一些实验结果。我们的演示展示了如何将 DiffusionGemma 模型适配为输出确定性概率,方法是暴露大语言模型生成的 logprobs。我们的初始方法建立在 Matt Mastracci 的独立研究基础之上,他一直在机器学习(ML)社区活跃,通过分享新想法和向 vLLM 推理引擎提交拉取请求,使 DiffusionGemma 获得更强的支持。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力