微软发布决策模型Microsoft-Decision-1:基于Qwen3.5
Microsoft AI Releases Microsoft-Decision-1: A Qwen3.5-9B Decision-Scoring Model
面向Agent控制的专用决策模型落地,低成本高吞吐特性值得开发者关注,建议评估其在自动化流程中的替换潜力。
Microsoft has released Microsoft-Decision-1, a decision model for routing, classification, verification and agent control. Microsoft-Decision-1 is a decision-scoring model that returns a calibrated probability for each fixed answer option instead of generated text. It is post-trained from Alibaba’s Qwen3.5-9B and available now in Microsoft Foundry and OpenRouter. .
微软发布了 Microsoft-Decision-1,这是一个用于路由、分类、验证和智能体控制的决策模型。Microsoft-Decision-1 是一个决策评分模型,它为每个固定的答案选项返回校准后的概率,而不是生成文本。它基于阿里巴巴的 Qwen3.5-9B 进行后训练,目前已在 Microsoft Foundry 和 OpenRouter 上提供。
TL;DR
TL;DR(太长不看版)
- Size: Built on Qwen3.5-9B; exact parameter count not disclosed. 32,768-token context window.
- Runs on: Hosted API only (Microsoft Foundry, OpenRouter via Azure). No open weights, no quantized variants, hardware not disclosed.
- Performance: Highest average accuracy in Microsoft’s 36-benchmark comparison, at 85 ms p50 latency.
- Best: 83.5% average accuracy across 36 benchmarks, ahead of Quyet-1.0-Large at 81.9%.
- Worst: Calibration of 92.2, second to Quyet-1.0-Large at 93.1. Text-only, no explanations.
- Bottom line:
- Best thing: fast, cheap, calibrated decisions at $0.042 per million input tokens with free output.
- Worst thing: closed weights, and every benchmark is vendor-run.
- 规模:基于 Qwen3.5-9B 构建;确切参数量未披露。支持 32,768 token 的上下文窗口。
- 运行环境:仅限托管 API(通过 Azure 在 Microsoft Foundry 和 OpenRouter 上提供)。无开源权重,无量化变体,硬件信息未披露。
- 性能:在微软的 36 项基准测试对比中,平均准确率最高,p50 延迟为 85 毫秒。
- 最佳表现:在 36 项基准测试中平均准确率为 83.5%,领先于 Quyet-1.0-Large 的 81.9%。
- 最差表现:校准度为 92.2,仅次于 Quyet-1.0-Large 的 93.1。仅支持文本输出,不提供解释。
- 结论:
- 最大优点:快速、廉价且校准良好的决策,每百万输入 token 仅需 0.042 美元,输出免费。
- 最大缺点:权重封闭,且所有基准测试均由厂商自行运行。
What is a decision model?
什么是决策模型?
A decision model reads an input and scores a closed set of options. It does not write prose. Microsoft frames decision models as a new AI category, built for outputs software can act on immediately.
决策模型读取输入并对一组封闭选项进行评分。它不撰写散文。微软将决策模型定义为一种新的 AI 类别,专为软件可立即执行的输出而构建。
How does Microsoft-Decision-1 work?
Microsoft-Decision-1 如何工作?
Microsoft post-trained Qwen3.5-9B for single-pass decision scoring. Given a situation, a question and fixed options, it returns a probability per option in one call. It plans to rebase future versions on MAI and OpenAI models.
微软对 Qwen3.5-9B 进行了单次传递决策评分的后训练。给定情境、问题和固定选项,它会在一次调用中返回每个选项的概率。微软计划将未来版本的基础重新建立在 MAI 和 OpenAI 模型之上。
The Foundry model card lists the supported formats:
Foundry 模型卡片列出了支持的格式:
- yes/no, multiple-choice, rating, classification and rubric questions
- grading of AI responses and proposed agent actions
- groundedness checks against supplied evidence
- explicit abstention options such as “cannot tell”
- 是/否、多项选择、评分、分类和量规问题
- 对 AI 响应和提议的智能体动作进行评分
- 针对提供的证据进行真实性检查
- 明确的弃权选项,如“无法判断”
Training used public datasets under Microsoft’s Open Data process plus synthetic data. Output is JSON. OpenRouter notes that weights update continually while the API shape stays fixed.
训练使用了微软开放数据流程下的公共数据集以及合成数据。输出为 JSON 格式。OpenRouter 指出,权重会持续更新,而 API 结构保持不变。
How fast and accurate is it?
它的速度和准确性如何?
Microsoft compared 9 systems across 36 benchmarks with 147,137 questions. Benchmarks were kept blind from training. Microsoft-Decision-1 led on average accuracy at 83.5%.
微软在 36 项基准测试中对 9 个系统进行了比较,共涉及 147,137 个问题。基准测试与训练过程保持隔离。Microsoft-Decision-1 以 83.5% 的平均准确率位居第一。
Its p50 latency was 85 ms, with p95 at 125 ms. That is 4.5 times quicker than Quyet-1.0-Large and 35 times quicker than GPT-6 Sol, which took 3.01 s.
其 p50 延迟为 85 毫秒,p95 延迟为 125 毫秒。这比 Quyet-1.0-Large 快 4.5 倍,比耗时 3.01 秒的 GPT-6 Sol 快 35 倍。
Microsoft also tested robustness. It perturbs each request 8 ways, including paraphrasing and option shuffling. The model flipped its decision on 1.3% of perturbations on average. Flips were zero when options were paraphrased, reversed or shuffled.
微软还测试了鲁棒性。它对每个请求进行8种方式的扰动,包括改写和选项重排。模型在平均1.3%的扰动情况下翻转了决策。当选项被改写、反转或重排时,翻转次数为零。
On safety, Microsoft ran 5,250 requests across 11 benchmarks. These covered harmful content, jailbreaks and prompt injection. Microsoft reports correct refusals with high retained utility, without publishing a score.
在安全性方面,微软在11个基准测试中运行了5,250个请求。这些测试涵盖了有害内容、越狱攻击和提示注入。微软报告称,模型在保持高实用性的同时正确拒绝了不当请求,但未公布具体分数。
Internal results Microsoft reports
微软报告的内部结果
- Xbox Research: sorted 10,000+ feedback items, 14 times faster and 200 times cheaper than GPT-6 Sol.
- Copilot quality control: competitive with GPT5.6 Luna and 100 times faster.
- Microsoft Discovery: 46 times more consistent than LLM scoring, at 3 times the speed.
- Xbox研究:对10,000多条反馈项进行了排序,速度比GPT-6 Sol快14倍,成本降低200倍。
- Copilot质量控制:与GPT5.6 Luna相当,且速度快100倍。
- 微软发现(Microsoft Discovery):一致性比LLM评分高出46倍,速度是后者的3倍。
How does it compare with other decision models?
它与其他决策模型相比如何?
| Feature | Microsoft-Decision-1 | Quyet-1.0-Large | H2O-Lightning-4B | GPT-6 Luna Decisions |
|---|---|---|---|---|
| Developer | Microsoft | Chinh Nguyen | H2O.ai | OpenAI |
| Base model | Qwen3.5-9B | Gemma-4-31B-it (LoRA merged) | Qwen3.5-4B | Not disclosed |
| Parameters | Not disclosed | 31.3B | 4B | Not disclosed |
| Context / input | 32,768 tokens | 8,000-token prompt (6,000 state) | 40,960 (vLLM serve setting) | Not disclosed |
| License | Proprietary API | Apache-2.0 | Apache-2.0 | Proprietary API |
| Modality | Text | Text | Text and images | Text and images |
| Hardware | Hosted (Azure) | 1x 80 GB GPU, bf16 | 32 GB GPU tested; 9.1 GB weights | Hosted |
| Accuracy (Microsoft, 36 benchmarks) | 83.5% | 81.9% | 77.2% | 79.4% |
| Calibration (Microsoft) | 92.2 | 93.1 | 91.8 | 89.9 |
| Median latency in Microsoft chart | 85 ms | 380 ms | 210 ms | 300 ms |
| Price | $0.042/M input, output free | Self-host | Self-host | $0.10/M input, output free |
| 特性 | Microsoft-Decision-1 | Quyet-1.0-Large | H2O-Lightning-4B | GPT-6 Luna Decisions |
|---|---|---|---|---|
| 开发者 | 微软 | Chinh Nguyen | H2O.ai | OpenAI |
| 基础模型 | Qwen3.5-9B | Gemma-4-31B-it (LoRA合并) | Qwen3.5-4B | 未公开 |
| 参数量 | 未公开 | 313亿 | 40亿 | 未公开 |
| 上下文/输入 | 32,768 tokens | 8,000-token提示词(6,000状态) | 40,960(vLLM服务设置) | 未公开 |
| 许可证 | 专有API | Apache-2.0 | Apache-2.0 | 专有API |
| 模态 | 文本 | 文本 | 文本和图片 | 文本和图片 |
| 硬件 | 托管(Azure) | 1x 80 GB GPU, bf16 | 测试用32 GB GPU;权重9.1 GB | 托管 |
| 准确率(微软,36个基准测试) | 83.5% | 81.9% | 77.2% | 79.4% |
| 校准度(微软) | 92.2 | 93.1 | 91.8 | 89.9 |
| 微软图表中的中位延迟 | 85 ms | 380 ms | 210 ms | 300 ms |
| 价格 | 输入$0.042/M,输出免费 | 自托管 | 自托管 | 输入$0.10/M,输出免费 |
Accuracy, calibration and latency rows come from Microsoft’s chart. Competitor latencies there use the JevBench v1.6.1 adjusted median.
准确率、校准度和延迟行数据来自微软的图表。其中的竞争对手延迟数据使用的是JevBench v1.6.1调整后的中位数。
Is the latency comparison apples to apples?
这种延迟比较是否公平可比?
Not fully. Microsoft measured its own model through Foundry. Competitor figures use JevBench’s adjusted median, not raw timings.
并不完全如此。微软通过Foundry测量了其自身模型的延迟。竞争对手的数据使用的是JevBench的调整后中位数,而非原始计时数据。
H2O.ai’s model card disputes this. It states that JevBench’s adjusted figure doubles measured time and adds 0.15 s. H2O says its model’s measured median is 29 ms, not 210 ms. Microsoft-Decision-1 does not appear on the JevBench board. On that board’s official composite, H2O-Lightning-4B ranks #1 at 72.5.
H2O.ai的模型卡片对此提出争议。它指出JevBench的调整后数据会使测量时间翻倍并增加0.15秒。H2O表示其模型的实际测量中位数为29毫秒,而非210毫秒。Microsoft-Decision-1并未出现在JevBench榜单上。在该榜单的官方综合排名中,H2O-Lightning-4B以72.5分位列第一。
How do developers deploy it?
开发者如何部署它?
Developers can call it from Microsoft Foundry as a generally available ‘Direct from Azure’ model. On OpenRouter, it runs on the Decisions API, not the chat endpoint. Chat completions SDKs will not work.
开发者可以将其作为一般可用的“Direct from Azure”模型在 Microsoft Foundry 中调用。在 OpenRouter 上,它运行于 Decisions API,而非聊天端点。聊天补全 SDK 将无法工作。
Pricing is $0.042 per million input tokens, with free output. OpenAI’s Luna decisions endpoint charges $0.10 per million input tokens.
定价为每百万输入令牌 0.042 美元,输出免费。OpenAI 的 Luna decisions 端点每百万输入令牌收费 0.10 美元。
What are the limitations?
有哪些限制?
The model card is explicit:
模型卡片明确说明:
- Not for text generation, open-ended Q&A, chat, translation or summarization.
- Text-only. No images, audio or video.
- No explanations or rationales in the output.
- Not for sole automated decisions on credit, employment, housing, healthcare or legal rights.
- Applications must define options, thresholds, escalation paths and human oversight.
- 不用于文本生成、开放式问答、聊天、翻译或摘要。
- 仅限文本。不支持图像、音频或视频。
- 输出中不包含解释或理由。
- 不用于信用、就业、住房、医疗保健或法律权利方面的单一自动化决策。
- 应用程序必须定义选项、阈值、升级路径和人工监督。
Key Takeaways
关键要点
- Microsoft-Decision-1 scores fixed options with calibrated probabilities, not text.
- Post-trained from Qwen3.5-9B, with a 32,768-token context window.
- Leads Microsoft’s 36-benchmark comparison at 83.5% accuracy and 85 ms p50.
- Costs $0.042 per million input tokens; output tokens are free.
- Latency comparisons are disputed; independent JevBench results are pending.
- Microsoft-Decision-1 对固定选项进行评分并给出校准后的概率,而非生成文本。
- 基于 Qwen3.5-9B 进行后训练,上下文窗口为 32,768 个令牌。
- 在微软的 36 项基准测试对比中领先,准确率为 83.5%,p50 延迟为 85 毫秒。
- 每百万输入令牌成本为 0.042 美元;输出令牌免费。
- 延迟比较存在争议;独立的 JevBench 结果尚待公布。
Check out the official announcement, the Foundry model card and the OpenRouter listing. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
请查看官方公告、Foundry 模型卡片以及 OpenRouter 上的列表。所有功劳归于该项目的研究人员。此外,欢迎在 Twitter 上关注我们,别忘了加入我们拥有 15 万+成员的 ML SubReddit,并订阅我们的新闻通讯。等等!你在 Telegram 上吗?现在你也可以加入我们的 Telegram 群组。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力