NVIDIA发布Nemotron 3.5 Lightning与NeMo
NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AI
As AI shifts from chatbots to autonomous agents, open models are serving market demands for full control over where AI runs and how it’s deployed and evolves.
随着人工智能从聊天机器人转向自主代理,开放模型正在满足市场对完全控制人工智能运行位置、部署方式及演进方式的需求。
Today, NVIDIA is expanding its Nemotron 3 model family with Nemotron 3.5 Lightning, the highest-efficiency model in its class for long-running agentic AI workloads. This release follows Nemotron 3 Nano and reflects NVIDIA’s commitment to continually improving open models for greater accuracy and speed.
今天,NVIDIA 正在扩展其 Nemotron 3 模型系列,推出 Nemotron 3.5 Lightning,这是同类中用于长期运行的代理式 AI 工作负载的最高效率模型。此次发布紧随 Nemotron 3 Nano 之后,体现了 NVIDIA 对持续改进开放模型以提高准确性和速度的承诺。
Built for specialized tasks within larger multi-agent systems, Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model, helps create smarter and more efficient agentic applications.
Nemotron 3.5 Lightning 专为大型多代理系统中的专门任务而构建,是一个 300 亿参数的专家混合模型,有助于创建更智能、更高效的代理式应用。
Also, NVIDIA is releasing NeMo Switchyard, an open source library for smart routing inside popular agent tools. Enterprises can use it to build a router based on their specific needs. When deployed, NeMo Switchyard can intelligently direct each request to the most capable and suitable model for the job without requiring developers to rewrite their applications.
此外,NVIDIA 还发布了 NeMo Switchyard,这是一个开源库,用于在流行的代理工具中进行智能路由。企业可以根据自身需求构建路由器。部署后,NeMo Switchyard 可以智能地将每个请求引导至最适合该任务的模型,而无需开发人员重写其应用程序。
Together, Nemotron 3.5 Lightning and NeMo Switchyard deliver greater control over how AI is deployed, where it runs and how efficiently it operates — across PCs, workstations, data centers and the cloud.
总之,Nemotron 3.5 Lightning 和 NeMo Switchyard 提供了对 AI 部署方式、运行位置以及运行效率的更大控制——涵盖 PC、工作站、数据中心和云端。
Nemotron 3.5 Lightning delivers frontier-level intelligence in a small, customizable open model built for high-volume agentic workflows.
Nemotron 3.5 Lightning 在一个小巧、可定制的开放模型中提供了前沿级别的智能,专为高容量代理工作流而构建。
Always-On Agents Need a System of Models
始终在线的代理需要模型系统
Modern agentic systems — always-on agents — increasingly operate as systems of models, or model ensembles, with different models specialized for different tasks.
现代代理系统——始终在线的代理——越来越多地作为模型系统或模型集成来运行,不同的模型专门用于不同的任务。
NVIDIA Nemotron open models are designed for this architecture. A frontier reasoning model such as Nemotron 3 Ultra or GPT-5.6 may plan and orchestrate a workflow, while smaller specialized models like Nemotron 3.5 Lightning can perform targeted tasks such as code review, tool use, security alert monitoring and answering billing questions.
NVIDIA Nemotron 开放模型专为此架构而设计。像 Nemotron 3 Ultra 或 GPT-5.6 这样的前沿推理模型可以规划并编排工作流,而像 Nemotron 3.5 Lightning 这样较小的专用模型可以执行针对性任务,如代码审查、工具使用、安全警报监控和回答计费问题。
Powering High-Volume Specialized Tasks With Nemotron 3.5 Lightning
使用 Nemotron 3.5 Lightning 驱动高容量专用任务
NVIDIA Nemotron 3.5 Lightning is a fully customizable open model built for high-volume tasks powering always-on agents. It was developed with contributions from the Nemotron Coalition, whose members provided evaluation methodologies, inference software and datasets to help advance the model.
NVIDIA Nemotron 3.5 Lightning 是一个完全可定制的开放模型,专为驱动始终在线代理的高容量任务而构建。它由 Nemotron 联盟的贡献开发而成,联盟成员提供了评估方法、推理软件和数据集,以帮助推进该模型。
The model delivers up to 4x faster output speed, leading to 30% faster agentic task completion compared with other models in its class. And because it’s open and customizable, Nemotron 3.5 Lightning can be easily post-trained with NVIDIA NeMo on an organization’s own domain data, tools and workflows to improve accuracy for specialized tasks.
该模型的输出速度最高提升至 4 倍,与同类其他模型相比,代理任务完成速度提升 30%。而且,由于它是开放且可定制的,Nemotron 3.5 Lightning 可以使用 NVIDIA NeMo 在组织自身的领域数据、工具和工作流程上进行轻松的后训练,以提高专业任务的准确性。
PinchBench benchmarks demonstrate that Nemotron 3.5 Lightning delivers faster agentic task completion with frontier-level accuracy compared to other models in its class.
PinchBench 基准测试表明,与同类其他模型相比,Nemotron 3.5 Lightning 在代理任务完成速度上更快,同时达到前沿水平的准确性。
AI leaders across industries are customizing Nemotron 3.5 Lightning for their workloads, including CrowdStrike for cybersecurity, Harvey with Trajectory for legal services and CodeRabbit with Baseten for code review, helping improve accuracy for domain-specific agentic tasks. Additionally, Lila Sciences is helping to improve reasoning capabilities for agentic tasks across physical and life sciences, and Fastino Labs customized the model and is seeing leading accuracies for software development, finance and healthcare workloads.
各行业的 AI 领导者正在为自身工作负载定制 Nemotron 3.5 Lightning,包括 CrowdStrike 用于网络安全,Harvey 与 Trajectory 合作用于法律服务,CodeRabbit 与 Baseten 合作用于代码审查,帮助提高特定领域代理任务的准确性。此外,Lila Sciences 正在帮助提升物理和生命科学领域代理任务的推理能力,Fastino Labs 定制了该模型,并在软件开发、金融和医疗保健工作负载方面取得了领先的准确性。
Enterprises have customized Nemotron 3.5 Lightning to achieve leading accuracy for their specialized task in their agentic workflows.
企业已定制 Nemotron 3.5 Lightning,以在其代理工作流中实现专业任务的领先准确性。
Nemotron 3.5 Lightning also gives organizations control over privacy and deployment. It can run on local AI systems — including NVIDIA RTX PCs, NVIDIA DGX Spark, NVIDIA DGX Station and NVIDIA Jetson — to help users maximize existing infrastructure investments, or scale across edge AI devices, NVIDIA RTX PRO workstations, data centers and cloud environments for enterprise use cases. And Nemotron 3.5 Lightning can run locally or on premises for high-volume, specialized tasks that require fast responses.
Nemotron 3.5 Lightning 还让组织能够控制隐私和部署。它可以在本地 AI 系统上运行,包括 NVIDIA RTX PC、NVIDIA DGX Spark、NVIDIA DGX Station 和 NVIDIA Jetson,帮助用户最大化现有基础设施投资,或跨边缘 AI 设备、NVIDIA RTX PRO 工作站、数据中心和云环境进行扩展,以满足企业用例。而且,Nemotron 3.5 Lightning 可以在本地或内部部署运行,适用于需要快速响应的高容量、专业任务。
Also, as with every Nemotron launch, NVIDIA publishes as much of the training data and techniques as licensing permits, which allows for traceability, auditing and training of other models. Alongside Lightning, NVIDIA is releasing Nemotron-RL-Agentic-Terminal-Pivot, an agentic reinforcement learning dataset used to post-train it for coding agent capabilities.
此外,与每次 Nemotron 发布一样,NVIDIA 在许可允许的范围内尽可能多地发布训练数据和技术,这有助于实现可追溯性、审计和其他模型的训练。与 Lightning 一同发布的还有 Nemotron-RL-Agentic-Terminal-Pivot,这是一个用于对其编码代理能力进行后训练的代理强化学习数据集。
More Efficient AI Apps With Model Routing
通过模型路由实现更高效的 AI 应用
Some models are better for coding, some for reasoning, some for lightweight tasks and some are optimized to run locally for greater privacy and efficiency. If customers rely on one default model, they might either overspend or lose quality; if they manage routing manually, it becomes integration work that can slow down a deployment.
有些模型更适合编码,有些更适合推理,有些适合轻量级任务,还有一些针对本地运行进行了优化,以提高隐私和效率。如果客户依赖一个默认模型,他们可能会过度支出或损失质量;如果他们手动管理路由,这就会变成集成工作,可能拖慢部署。
NVIDIA NeMo Switchyard is an open source model routing library for AI agents. The technology routes prompts to the most capable and efficient model for each step of an agent workflow automatically, based on specific needs. Agent application developers can tune or modify the router with different routing algorithms to match their priorities, such as quality, latency and cost requirements. In a system of models, enterprises can create powerful AI agents with improved tokenomics.
NVIDIA NeMo Switchyard 是一个面向 AI 代理的开源模型路由库。该技术可根据具体需求,自动将提示词路由到代理工作流每一步中最有能力且最高效的模型。代理应用开发者可以使用不同的路由算法调整或修改路由器,以匹配其优先级,例如质量、延迟和成本要求。在模型系统中,企业可以创建具有改进的 token 经济性的强大 AI 代理。
Internal benchmarks show that NeMo Switchyard maintains frontier-level accuracy while reducing task completion cost to nearly one-third of Opus 4.8 alone.
内部基准测试显示,NeMo Switchyard 在保持前沿级准确性的同时,将任务完成成本降低到仅使用 Opus 4.8 时的近三分之一。
NVIDIA internal benchmarks show that NeMo Switchyard maintains frontier-level accuracy while reducing task completion cost to nearly one-third of Opus 4.8 alone.
NVIDIA 内部基准测试显示,NeMo Switchyard 在保持前沿级准确性的同时,将任务完成成本降低到仅使用 Opus 4.8 时的近三分之一。
NVIDIA is working with partners across the AI ecosystem to bring intelligent model routing into the tools and platforms developers already use.
NVIDIA 正在与 AI 生态系统中的合作伙伴合作,将智能模型路由引入开发者已经使用的工具和平台。
- Boomi: Evaluated Switchyard across five routing capabilities, achieving 100% domain-routing accuracy, sending 59% of traffic to a 5x faster fine-tuned model and reducing later-turn latency by 21%.
- Cadence: Improved efficiency by 9.9% by using the ChipStack AI Super Agent for a formal verification use case.
- Classmethod: Is running opencode and Fireworks workloads using NeMo Switchyard internally, with initial testing showing a 27% cost reduction while maintaining quality.
- Cognition: Integrated the NVIDIA NeMo Switchyard staged router into Devin Desktop for NVIDIA internal use, achieving near-frontier performance on FrontierCode Main while reducing mean cost by 28% relative to routing all requests to a single underlying frontier model.
- Kong: Delivers routing with NeMo Switchyard natively through Kong AI Gateway.
- LangChain: With NeMo Switchyard, achieved 74% lower cost in 145 multi-turn Deep Agents tasks by routing only 7% of calls to a frontier model, at a 6% accuracy tradeoff.
- LiteLLM: Is adding NeMo Switchyard as a plug-in into its proxy layer so developers can access these benefits without changing their existing stack.
- Nous Research: Integrated NeMo Switchyard into Hermes to provide developers with an easy-to-configure routing system to improve agent efficiency.
- Ramp: Used NeMo Switchyard to match a frontier model’s performance while cutting costs by 58% and runtime by 33% in Ramp SWE-Bench.
- Siemens: Is benchmarking to improve efficiency in its Fuse EDA AI Agent.
- Boomi:在五项路由能力上评估了 Switchyard,实现了 100% 的领域路由准确性,将 59% 的流量发送到速度快 5 倍的微调模型,并将后期轮次延迟降低了 21%。
- Cadence:通过将 ChipStack AI 超级代理用于形式验证用例,效率提高了 9.9%。
- Classmethod:在内部使用 NeMo Switchyard 运行 opencode 和 Fireworks 工作负载,初步测试显示成本降低了 27%,同时保持了质量。
- Cognition:将 NVIDIA NeMo Switchyard 分阶段路由器集成到 Devin Desktop 中供 NVIDIA 内部使用,在 FrontierCode Main 上实现了接近前沿的性能,同时相对于将所有请求路由到单个底层前沿模型,平均成本降低了 28%。
- Kong:通过 Kong AI Gateway 原生提供 NeMo Switchyard 路由。
- LangChain:使用 NeMo Switchyard,在 145 个多轮 Deep Agents 任务中,通过仅将 7% 的调用路由到前沿模型,成本降低了 74%,同时准确性损失了 6%。
- LiteLLM:正在将 NeMo Switchyard 作为插件添加到其代理层中,以便开发者无需更改现有技术栈即可获得这些好处。
- Nous Research:将 NeMo Switchyard 集成到 Hermes 中,为开发者提供易于配置的路由系统,以提高代理效率。
- Ramp:在 Ramp SWE-Bench 中使用 NeMo Switchyard 匹配了前沿模型的性能,同时成本降低了 58%,运行时间缩短了 33%。
- Siemens:正在对其 Fuse EDA AI 代理进行基准测试,以提高效率。
Nemotron 3.5 Lightning is available on Hugging Face, ModelScope, OpenRouter and build.nvidia.com as an NVIDIA NIM microservice as well as through a broad ecosystem of NVIDIA Cloud Partners, post-training platforms, inference platforms and cloud service providers. NeMo Switchyard is available on GitHub and coming to partner platforms soon.
Nemotron 3.5 Lightning 已在 Hugging Face、ModelScope、OpenRouter 和 build.nvidia.com 上以 NVIDIA NIM 微服务的形式提供,同时也可通过 NVIDIA 云合作伙伴、后训练平台、推理平台和云服务提供商的广泛生态系统获取。NeMo Switchyard 已在 GitHub 上发布,并即将登陆合作伙伴平台。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力