Cloudflare发布Clef:基于Qwen3.8-27B的多模态决策模型
Clef: Open Weights decision model by Cloudflare
基于Qwen底座的多模态决策模型是Agent落地的关键基础设施,Clef提供了完整的权重、代码和API兼容性,做业务逻辑自动化的同学值得研究。
Clef
- Announcement: Clef decision models on the Cloudflare blog
- Decision Index leaderboard: clef-evals.workers-ai-mle.workers.dev
- 公告:Cloudflare 博客上的 Clef 决策模型
- 决策索引排行榜:clef-evals.workers-ai-mle.workers.dev
Clef is a 27B multimodal model that turns a state and a schema of typed questions into decisions. It reads the state as text, JSON, images, or video, and returns a probability for every allowed option of every question in a single forward pass. There is no free-form text generation and no output parsing.
Clef 是一个 27B 参数的多模态模型,它将状态和一组类型化问题的模式转化为决策。它以文本、JSON、图像或视频的形式读取状态,并在单次前向传播中为每个问题的所有允许选项返回概率。不支持自由形式的文本生成,也不进行输出解析。
The Clef API is fully compatible with Jev and SystemOne.
Clef API 与 Jev 和 SystemOne 完全兼容。
Clef is post-trained from Qwen/Qwen3.8-27B. See Clef-Flash for the smaller, faster variant.
Clef 是在 Qwen/Qwen3.8-27B 基础上进行后训练得到的。请参阅 Clef-Flash 以获取更小、更快的变体。
Model
模型
- Backbone: Qwen/Qwen3.8-27B with its vision encoder, stored as standard sharded safetensors.
- Joint schema head: a small transformer head that reads the backbone's final hidden states, routes evidence from the state to each question, and scores all options of all questions jointly.
- Output: one logit per allowed option for each question. Apply a softmax per question to get probabilities.
- 主干网络:Qwen/Qwen3.8-27B 及其视觉编码器,存储为标准分片 safetensors 格式。
- 联合模式头(Joint schema head):一个小型 Transformer 头,用于读取主干网络的最终隐藏状态,将来自状态的证据路由到每个问题,并对所有问题的所有选项进行联合评分。
- 输出:每个问题的每个允许选项对应一个 logit。对每个问题应用 softmax 以获得概率。
Files
文件
| File | Purpose |
|---|---|
| model-*.safetensors, model.safetensors.index.json, config.json, generation_config.json | Backbone, including the vision encoder |
| joint_head.safetensors, joint_head_config.json | Joint schema head |
| joint_schema_model.py | Record encoding, batching, the model, load_release_model, and systemone |
| tokenizer.json, tokenizer_config.json, chat_template.jinja, processor_config.json | Tokenizer and image/video processor |
| LICENSE | Apache-2.0 license |
| 文件 | 用途 |
|---|---|
| model-*.safetensors, model.safetensors.index.json, config.json, generation_config.json | 主干网络,包括视觉编码器 |
| joint_head.safetensors, joint_head_config.json | 联合模式头 |
| joint_schema_model.py | 记录编码、批处理、模型、load_release_model 以及 systemone |
| tokenizer.json, tokenizer_config.json, chat_template.jinja, processor_config.json | 词元器和图像/视频处理器 |
| LICENSE | Apache-2.0 许可证 |
Usage
用法
Tested with torch 2.11 and transformers 5.10.2 on a single H200. Image and video inputs also need pillow.
在单张 H200 GPU 上使用 torch 2.11 和 transformers 5.10.2 进行了测试。图像和视频输入还需要 pillow。
import sys
import torch
from huggingface_hub import snapshot_download
path = snapshot_download("Cloudflare/clef")
sys.path.insert(0, path)
from joint_schema_model import collate_records, encode_record, load_release_model
model, processor = load_release_model(path, device="cuda")
record = {
"state": {"invoice": {"vendor": "Acme", "total": 1250.0, "currency": "USD", "status": "overdue"}},
"questions": {
"status": {
"type": "choice",
"instructions": "What is the invoice status?",
"criteria": {"paid": "Invoice is paid.", "overdue": "Invoice is past due.", "draft": "Not sent."},
},
"large": {"type": "noul", "instructions": "Is the total above 1000 USD?"},
},
}
encoded = encode_record(processor.tokenizer, record, processor=processor)
batch = collate_records([encoded], processor.tokenizer.pad_token_id, torch.device("cuda"))
with torch.inference_mode():
logits = model(batch)[0]
for question, question_logits in zip(encoded.questions, logits):
probabilities = question_logits.float().softmax(-1).tolist()
print(question.question_id, dict(zip(question.option_ids, probabilities)))Jev / SystemOne API
systemone takes a Jev/SystemOne POST /v1/systemone request body and returns the same response body: model, answers keyed by question ID, and usage. A choice answer has choice, confidence, and probabilities; a score answer has the expected score, confidence, legend, and probabilities; a noul answer has the probability of true. instructions is optional, and images and videos may be added to the request.
systemone 接收 Jev/SystemOne POST /v1/systemone 请求体,并返回相同的响应体:包含模型、按问题 ID 键值排列的答案以及使用情况。choice 类型的回答包含 choice、confidence 和 probabilities;score 类型的回答包含预期分数、confidence、legend 和 probabilities;noul 类型的回答包含 true 的概率。instructions 是可选的,并且可以向请求中添加图像和视频。
from joint_schema_model import systemone
response = systemone(model, processor, {
"model": "clef",
"state": "Our checkout started returning errors and orders are blocked.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle the message?",
"criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
},
"urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
"outage": {"type": "noul", "instructions": "Is a service down?"},
},
})
print(response["answers"])Images and video
图像和视频
Add images (PIL images) or videos (frame arrays) to the record and pass the processor to encode_record. Optional processor arguments go in media_kwargs.
将图像(PIL 图像)或视频(帧数组)添加到记录中,并将处理器传递给 encode_record。可选的处理器参数放在 media_kwargs 中。
from PIL import Image
record = {
"state": {"task": "Review the attached receipt."},
"images": [Image.open("receipt.jpg")],
"questions": {
"legible": {"type": "noul", "instructions": "Is the receipt total legible?"},
},
}
encoded = encode_record(processor.tokenizer, record, processor=processor)Text-only and multimodal records can be mixed in the same batch.
纯文本记录和 multimodal 记录可以在同一个批次中混合使用。
Input format
输入格式
| Field | Description |
|---|---|
| state | Any string or JSON value describing the situation to decide on |
| images, videos | Optional lists of images or video frame arrays |
| media_kwargs | Optional keyword arguments for the image/video processor |
| questions | Mapping of question ID to question |
| 字段 | 描述 |
|---|---|
| state | 描述待决策情况的任意字符串或 JSON 值 |
| images, videos | 可选的图像列表或视频帧数组 |
| media_kwargs | 图像/视频处理器的可选关键字参数 |
| questions | 问题 ID 到问题的映射 |
Each question has:
每个问题包含:
- type: noul (true/false), choice (named options), or score (ordered options)
- instructions: what to decide; optional, and the question ID is used when it is omitted
- criteria: for choice, a mapping of option ID to description; for score, a list of option descriptions indexed from 0; for noul, optional descriptions for true and false
- 类型:noul(布尔值)、choice(命名选项)或 score(有序选项)
- 指令:需要做出的决策;可选,若省略则使用问题 ID
- 标准:对于 choice 类型,为选项 ID 到描述的映射;对于 score 类型,为从 0 开始索引的选项描述列表;对于 noul 类型,为 true 和 false 的可选描述
encode_record accepts max_length (default 16,384 tokens) and max_state_tokens to bound the input.
encode_record 接受 max_length(默认 16,384 个 token)和 max_state_tokens 以限制输入大小。
Results
结果
Decision Index
决策指数
Per-benchmark results from our internal run of the Decision Index 0.2.1 suite. Scores are percentages; ForecastBench is a Brier score, where lower is better. The last two rows are request latency in milliseconds, where lower is better. The best value in each row is in bold.
来自我们内部运行 Decision Index 0.2.1 套件的性能基准测试结果。分数为百分比;ForecastBench 采用 Brier 分数,数值越低越好。最后两行为请求延迟(毫秒),数值越低越好。每行中的最佳值以粗体显示。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力