NVIDIA 发布 Cosmos 3 开放世界模型,推动物理 AI 发展
Into the Omniverse: How Open World Models Push the Frontier of Physical AI
做机器人、自动驾驶和视觉 AI 的同学注意了,NVIDIA 的 Cosmos 3 把世界模型做成了开放权重全家桶,还拿了多个榜单第一,赶紧去 Hugging Face 上看看能不能直接后训练。
Editor’s note: This post is part of Into the Omniverse, a series focused on how developers, 3D practitioners and enterprises can transform their workflows using the latest advancements in OpenUSD and NVIDIA Omniverse.
编者按:本文是“进入 Omniverse”系列文章之一,该系列聚焦开发者、3D 从业者及企业如何利用 OpenUSD 和 NVIDIA Omniverse 的最新进展来转变其工作流程。
In July, NVIDIA joined more than 200 companies and organizations in signing “Open Weights and American AI Leadership,” an open letter arguing that AI leadership will be measured not by any single frontier model but by whether an open ecosystem reaches every sector.
7 月,NVIDIA 与 200 多家公司和组织共同签署了“开放权重与美国 AI 领导力”公开信,该信主张 AI 领导力并非由任何单一前沿模型来衡量,而是取决于开放生态系统是否覆盖每个行业。
Open models, which anyone can download, inspect, modify and run on their own infrastructure, are what make that possible. Nowhere is that more crucial than in physical AI, where every deployment is a specialization problem.
开放模型允许任何人下载、检查、修改并在自己的基础设施上运行,正是这些模型使上述愿景成为可能。这一点在物理 AI 领域尤为关键,因为每一次部署都是一个专门化问题。
Physical AI has to understand and predict consequences, not just appearances.
物理 AI 必须理解并预测后果,而不仅仅是表象。
To make this possible, world models learn how physical environments behave, what may happen next and which following actions make sense. They can generate physically grounded world and action data, simulate future states and provide a foundation that teams can specialize for a robot, autonomous vehicle or vision AI system.
为了实现这一点,世界模型学习物理环境如何运作、接下来可能发生什么以及哪些后续动作是合理的。它们可以生成基于物理的世界和动作数据,模拟未来状态,并为团队提供基础,以便针对机器人、自动驾驶汽车或视觉 AI 系统进行专门化。
Open world models are already being used to generate training data, test policies and specialize physical AI systems. NVIDIA Cosmos 3 brings these capabilities together in an open model family, with leading benchmark results and adoption across robotics, autonomous vehicles and vision AI.
开放世界模型已被用于生成训练数据、测试策略以及专门化物理 AI 系统。NVIDIA Cosmos 3 在一个开放模型系列中整合了这些能力,在机器人、自动驾驶汽车和视觉 AI 领域取得了领先的基准测试结果并得到广泛采用。
And NVIDIA Omniverse libraries, part of NVIDIA Agent Toolkit, provides prebuilt capabilities for building simulation-ready worlds that physical AI teams can use to train, test and validate systems before real-world deployment.
作为 NVIDIA Agent Toolkit 的一部分,NVIDIA Omniverse 库提供了用于构建模拟就绪世界的预构建能力,物理 AI 团队可在实际部署前使用这些能力来训练、测试和验证系统。
World Models Are the Foundation of Physical AI
世界模型是物理 AI 的基础
The data behind physical AI is difficult and expensive to collect at the scale required. Rare events and long-tail scenarios can be especially difficult to reproduce safely and repeatedly.
物理 AI 所需的数据在所需规模下难以且昂贵地收集。罕见事件和长尾场景尤其难以安全且可重复地复现。
World models enable:
世界模型能够实现:
- More useful data by learning physical relationships from large-scale multimodal scenarios.
- More diverse environments that vary in weather, lighting, objects and trajectories.
- A better foundation to build on and adapt to a particular robot, vehicle, sensor configuration, task or operating environment.
- 通过从大规模多模态场景中学习物理关系,获得更有用的数据。
- 更多样化的环境,涵盖不同的天气、光照、物体和轨迹。
- 更好的基础,以便针对特定机器人、车辆、传感器配置、任务或运行环境进行构建和调整。
https://blogs.nvidia.com/wp-content/uploads/2026/08/cosmos-corp-blog-promo-1920x1080-1.mp4
https://blogs.nvidia.com/wp-content/uploads/2026/08/cosmos-corp-blog-promo-1920x1080-1.mp4
A general model hasn’t seen a team’s particular robot, sensors or operating environment. Closing that gap requires access to model weights, a license that permits adaptation and the tools needed for post-training.
通用模型尚未见过团队特定的机器人、传感器或运行环境。要缩小这一差距,需要访问模型权重、允许修改的许可证以及用于后训练的工具。
NVIDIA Cosmos world foundation models are available under the Linux Foundation’s OpenMDW 1.1 license, enabling teams to post-train models on their own data and hardware. Specialization is where openness becomes a practical technical requirement.
NVIDIA Cosmos 世界基础模型在 Linux 基金会的 OpenMDW 1.1 许可证下提供,使团队能够使用自己的数据和硬件对模型进行后训练。专业化是开放性成为实际技术需求的地方。
Specializing a model is only part of the workflow. Teams also need environments to generate data, run simulations and test behavior.
专业化模型只是工作流程的一部分。团队还需要环境来生成数据、运行模拟和测试行为。
Omniverse libraries help developers build simulation-ready environments, while OpenUSD provides the open framework for composing, reusing and exchanging complex 3D data across digital twins, simulations and synthetic data generation workflows. Together, Omniverse and OpenUSD cut the duplicated work that can otherwise pile up every time assets, sensor configurations or environmental conditions change.
Omniverse 库帮助开发者构建模拟就绪的环境,而 OpenUSD 提供了开放框架,用于在数字孪生、模拟和合成数据生成工作流程中组合、重用和交换复杂的 3D 数据。Omniverse 和 OpenUSD 共同减少了每当资产、传感器配置或环境条件变化时可能累积的重复工作。
Cosmos 3: The Frontier Model
Cosmos 3:前沿模型
NVIDIA Cosmos 3 — a frontier open physical AI foundation omni-model built on a mixture-of-transformers architecture — combines vision reasoning, world generation and action prediction, letting developers use one model family to understand scenes, generate synthetic data, simulate future states and build specialized world action models.
NVIDIA Cosmos 3——一个基于混合变换器架构的前沿开放物理 AI 全能模型——结合了视觉推理、世界生成和动作预测,让开发者可以使用一个模型系列来理解场景、生成合成数据、模拟未来状态并构建专门的世界动作模型。
Developers can use Cosmos 3 as a vision language model, as a physics-grounded world simulator that predicts future world states and generates large-scale synthetic data, or as the backbone for world action models, instead of assembling and maintaining a separate model for each capability.
开发者可以将 Cosmos 3 用作视觉语言模型、作为预测未来世界状态并生成大规模合成数据的物理基础世界模拟器,或作为世界动作模型的骨干,而不是为每种能力组装和维护单独的模型。
The family includes Cosmos 3 Super (64B) for high-fidelity world modeling, Cosmos 3 Nano (16B) for efficient reasoning and post-training, and Cosmos 3 Edge (4B) for on-device vision reasoning and robot policy deployment. Lightweight enough to run on edge GPUs, Cosmos 3 Edge can be deployed across NVIDIA RTX GPUs, NVIDIA DGX systems and NVIDIA Jetson, including Jetson Thor platforms.
该系列包括用于高保真世界建模的 Cosmos 3 Super (64B)、用于高效推理和后训练的 Cosmos 3 Nano (16B),以及用于设备端视觉推理和机器人策略部署的 Cosmos 3 Edge (4B)。Cosmos 3 Edge 足够轻量,可以在边缘 GPU 上运行,可部署在 NVIDIA RTX GPU、NVIDIA DGX 系统和 NVIDIA Jetson(包括 Jetson Thor 平台)上。
Across benchmark evaluations, Cosmos 3 ranks No. 1 on Artificial Analysis for open weights text-to-image and image-to-video generation, on PAI-Bench for world generation and in the image-to-video category of Physics-IQ. For robot policy, it ranks No. 1 on RoboLab. Cosmos 3 Super is also the highest-ranked open model on VANTAGE-Bench for vision understanding.
在基准评估中,Cosmos 3 在 Artificial Analysis 的开放权重文本到图像和图像到视频生成中排名第一,在 PAI-Bench 的世界生成和 Physics-IQ 的图像到视频类别中排名第一。对于机器人策略,它在 RoboLab 上排名第一。Cosmos 3 Super 也是 VANTAGE-Bench 视觉理解中排名最高的开放模型。
In addition to Cosmos, NVIDIA’s physical AI stack includes Isaac GR00T for robotics, Alpamayo for autonomous vehicles and Metropolis for vision AI.
除了 Cosmos,NVIDIA 的物理 AI 堆栈还包括用于机器人的 Isaac GR00T、用于自动驾驶汽车的 Alpamayo 和用于视觉 AI 的 Metropolis。
How Developers Are Putting Cosmos 3 to Work
开发者如何应用 Cosmos 3
Across industries, developers are building on NVIDIA Cosmos for physical AI applications: Doosan Robotics, LG Electronics, Samsung Electronics and Skild AI in robotics; Li Auto, Xiaomi and Afari in autonomous vehicles; and Centific, Fogsphere, Linker Vision, Milestone Systems and Yuan for vision AI agents powering industrial AI and smart spaces applications.
在各行各业,开发者们正基于NVIDIA Cosmos构建物理AI应用:机器人领域的Doosan Robotics、LG Electronics、Samsung Electronics和Skild AI;自动驾驶领域的Li Auto、Xiaomi和Afari;以及视觉AI代理领域的Centific、Fogsphere、Linker Vision、Milestone Systems和Yuan,这些代理正为工业AI和智能空间应用提供动力。
The NVIDIA Cosmos Coalition extends this work by bringing together world model builders, AI developers and physical AI leaders to contribute models, research and evaluation methods. NVIDIA recently expanded the coalition to Japan, where robotics and manufacturing leaders intend to join and develop open world models for factories, logistics, agriculture, construction, healthcare and transportation.
NVIDIA Cosmos联盟通过汇集世界模型构建者、AI开发者和物理AI领导者,共同贡献模型、研究和评估方法,扩展了这项工作。NVIDIA最近将该联盟扩展至日本,那里的机器人和制造业领导者有意加入,并为工厂、物流、农业、建筑、医疗保健和交通运输开发开放世界模型。
Together, these implementations and collaborations are establishing open world models as an adaptable foundation for physical AI across robots, autonomous vehicles and vision AI systems.
这些实施和合作共同将开放世界模型确立为物理AI在机器人、自动驾驶汽车和视觉AI系统中的适应性基础。
Get Plugged In
深入了解
Learn more about world models, OpenUSD and physical AI development by exploring these resources:
通过探索以下资源,了解更多关于世界模型、OpenUSD和物理AI开发的信息:
- Explore the open Cosmos 3 model collection and datasets on Hugging Face and GitHub.
- Read the Cosmos 3 technical report for full architecture details and evaluations.
- Read the Cosmos 3 technical blog.
- Tune in to the Cosmos Labs livestreams.
- Learn about the NVIDIA Cosmos Coalition.
- 在Hugging Face和GitHub上探索开放的Cosmos 3模型集合和数据集。
- 阅读Cosmos 3技术报告,了解完整的架构细节和评估。
- 阅读Cosmos 3技术博客。
- 观看Cosmos Labs直播。
- 了解NVIDIA Cosmos联盟。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力