跳到主内容
@wquguru
精选88NVIDIA 博客(RSS)产品发布/更新多源精选 ×2

NVIDIA IFA 2026:发布 RTX Spark PC 与 PAIR

Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026

原文
发到 X
推荐理由

本地 AI 部署门槛大幅降低,RTX Spark 硬件与 PAIR 分布式算力方案值得关注,开发者可参考其多模型适配清单优化本地链路。

Frontier intelligence is going local. At IFA 2026, NVIDIA, Microsoft and its partners are teaming up to provide faster inference and new tools that make agents easier to set up and run locally on NVIDIA hardware. New compact NVIDIA RTX Spark Windows PCs are also coming in October to give AI enthusiasts, developers and creators more ways to run capable agents locally and securely.

前沿智能正在走向本地化。在 2026 年国际消费电子展(IFA)上,英伟达、微软及其合作伙伴正携手合作,提供更快的推理速度和新工具,使代理能够在英伟达硬件上更轻松地设置并在本地运行。全新的紧凑型英伟达 RTX Spark Windows 电脑也将于十月上市,为 AI 爱好者、开发者和创作者提供更多在本地安全运行高性能代理的方式。

Today’s announcements include:

今天的公告包括:

  • Simplified local AI support for NVIDIA GPUs is coming in Hermes Agent, OpenClaw and Perplexity Portable Computer.
  • Up to 1.9x faster local inference — new llama.cpp and vLLM optimizations are available now directly and through LM Studio and Ollama.
  • NVIDIA PAIR — a Personal AI Router tool that intelligently distributes AI inference across the PCs on a user’s local network.
  • NVIDIA RTX Spark arrives in October — with new Windows PCs from Lenovo and Acer. Electronic Arts, Embark and Ubisoft are among the latest game publishers and developers bringing their blockbuster titles to NVIDIA RTX Spark.
  • Hermes Agent、OpenClaw 和 Perplexity Portable Computer 即将推出针对英伟达 GPU 的简化本地 AI 支持。
  • 本地推理速度最高提升 1.9 倍——新的 llama.cpp 和 vLLM 优化现已直接可用,也可通过 LM Studio 和 Ollama 获取。
  • NVIDIA PAIR ——一款个人 AI 路由工具,可智能地在用户本地网络中的各台 PC 之间分配 AI 推理任务。
  • NVIDIA RTX Spark 将于十月到来——联想和宏碁将推出新款 Windows 电脑。Electronic Arts、Embark 和育碧等最新的游戏发行商和开发商正将其重磅作品带入 NVIDIA RTX Spark。

Also, August was a busy month for local AI:

此外,八月也是本地 AI 繁忙的一个月:

  • Nemotron 3.5 Lightning — which can run on NVIDIA RTX PCs, RTX PRO Workstations, DGX Spark and Jetson — is a 30-billion parameter model that has been launched. Get started with Nemotron 3.5 Lightning today.
  • Z.ai’s GLM-5.3-Flash is a multimodal mixture-of-experts (MoE) model that’s bringing agentic AI to DGX Station.
  • Qwen has released Qwen3.8-Flash-Next, an open weight multimodal MoE model, which can run locally on DGX Spark and DGX Station, along with Qwen3.8-27B, a 27-billion-parameter open model optimized for local agentic and coding workloads on NVIDIA GPUs.
  • LTX’s LTX 2.5 is an open-world video generation model optimized for NVIDIA RTX GPUs, DGX Spark and DGX Station, with new NVFP4, FastVideo and ComfyUI enhancements for faster, more memory-efficient local generation.
  • MiniMax-H3 is an open-weight video generation model with synchronized audio that can run locally on NVIDIA GPUs through ComfyUI. FastVideo teamed up with NVIDIA researchers to improve this further by releasing FastH3 — an open-weight, four-step distilled version that improves performance by 7x. Optimized FastVideo recipes for NVIDIA RTX GPUs and DGX Spark are coming soon.
  • Meta’s Muse Glimmer is a 30-billion-parameter open-weight model for coding and agentic workloads that can run locally on GeForce RTX PCs, DGX Spark, DGX Station and Jetson. NVIDIA has also released NVFP4 quantization with DGX Spark support for more memory-efficient local deployment.
  • DeepSeek v4 Flash is a 284-billion-parameter MoE model with 13 billion active parameters that can run locally on 2x DGX Spark cluster and DGX Station.
  • Nemotron 3.5 Lightning ——可在英伟达 RTX PC、RTX PRO 工作站、DGX Spark 和 Jetson 上运行的 300 亿参数模型已发布。立即开始使用 Nemotron 3.5 Lightning。
  • Z.ai 的 GLM-5.3-Flash 是一个多模态混合专家(MoE)模型,正将代理式 AI 引入 DGX Station。
  • Qwen 发布了 Qwen3.8-Flash-Next,这是一个开源权重的多模态 MoE 模型,可在 DGX Spark 和 DGX Station 上本地运行;同时发布的还有 Qwen3.8-27B,这是一个 270 亿参数的开源模型,针对英伟达 GPU 上的本地代理和编码工作负载进行了优化。
  • LTX 的 LTX 2.5 是一个开放世界视频生成模型,针对英伟达 RTX GPU、DGX Spark 和 DGX Station 进行了优化,并引入了新的 NVFP4、FastVideo 和 ComfyUI 增强功能,以实现更快、更节省内存的本地生成。
  • MiniMax-H3 是一个带有同步音频的开源权重视频生成模型,可通过 ComfyUI 在英伟达 GPU 上本地运行。FastVideo 与英伟达研究人员合作,通过发布 FastH3——一个开源权重的四步蒸馏版本,将性能提升了 7 倍。针对英伟达 RTX GPU 和 DGX Spark 优化的 FastVideo 配方即将推出。
  • Meta 的 Muse Glimmer 是一个用于编码和代理工作负载的 300 亿参数开源权重模型,可在 GeForce RTX PC、DGX Spark、DGX Station 和 Jetson 上本地运行。英伟达还发布了支持 DGX Spark 的 NVFP4 量化技术,以实现更节省内存的本地部署。
  • DeepSeek v4 Flash 是一个拥有 2840 亿参数的 MoE(混合专家)模型,其中 130 亿为活跃参数,可在 2x DGX Spark 集群和 DGX Station 上本地运行。

A Simpler Start for Local Agents

本地智能体更简单的入门方式

Getting a local agent up and running with local models required some effort — choosing a model, finding a compatible inference server, dialing in quantization settings and keeping everything updated. That friction is disappearing on RTX and DGX systems.

要让本地智能体配合本地模型运行起来需要付出一些努力——选择模型、寻找兼容的推理服务器、调整量化设置并保持一切更新。这种摩擦正在 RTX 和 DGX 系统上消失。

Three of the most widely used agent apps will offer simplified local model setup on Windows, each built on llama.cpp and incorporating NVIDIA’s latest inference optimizations. The new setup experiences are designed to reduce manual configuration and make it easier to get local agents up and running.

三款使用最广泛的智能体应用将在 Windows 上提供简化的本地模型设置,它们均基于 llama.cpp 构建,并融入了 NVIDIA 最新的推理优化。新的设置体验旨在减少手动配置,使本地智能体的部署和运行更加轻松。

Last month, Perplexity introduced its Portable Computer agent, giving users a simple way to run Perplexity locally on Linux systems like NVIDIA DGX Spark with the models, orchestration and tools packaged into a single app experience.

上个月,Perplexity 推出了其 Portable Computer 智能体,为用户提供了一种在 Linux 系统(如 NVIDIA DGX Spark)上本地运行 Perplexity 的简便方法,将模型、编排和工具打包成单一的应用体验。

Perplexity Portable Computer is available on NVIDIA RTX GPUs with at least 24GB VRAM running on Linux, with support on Windows coming soon, bringing that same streamlined setup to a broader group of PC users. Users can run complete workflows locally without consuming credits, while selectively escalating parts of a task to one of 15+ frontier models in the cloud when additional research or reasoning is needed. Portable Computer asks for permission before sending content to the cloud, helping users keep sensitive information on their device. Here’s some example use-cases:

Perplexity Portable Computer 适用于配备至少 24GB VRAM 的 NVIDIA RTX GPU 并在 Linux 上运行的设备,Windows 支持即将推出,从而将该 streamlined(精简)的设置体验带给更广泛的 PC 用户群体。用户可以在本地运行完整的工作流程而无需消耗积分,同时在需要额外研究或推理时,有选择地将任务的部分内容升级到云端 15 多种前沿模型之一。Portable Computer 在向云端发送内容之前会征求用户许可,帮助用户将敏感信息保留在设备上。以下是一些示例用例:

  • Engineering: Review open PRs in a connected GitHub repo and sort them into ready, blocked, stale, and needs review, each tagged with the next step. Docs that fell out of sync with the latest merge get caught and fixed, with a PR opened for the changes.
  • Finance: Point the agent at two years of brokerage summaries, consolidated 1099s and tax returns and have it trace the recurring holdings creating the most avoidable fees and tax drag, with every figure cited to the exact file and page — all without a document ever reaching a chatbot.
  • Startups: Ask why activation went flat, and the agent analyzes the funnel export locally to find where new signups drop off between install and first completed task, then posts the top insights straight to the team’s Slack channel.
  • 工程:审查已连接的 GitHub 仓库中的开放 PR(拉取请求),并将它们分类为 ready(就绪)、blocked(阻塞)、stale(陈旧)和 needs review(需审核),每个类别都标记了下一步操作。那些与最新合并版本不同步的文档会被捕获并修复,同时为这些更改打开一个 PR。
  • 财务:将智能体指向两年的经纪商摘要、汇总的 1099 表格和报税表,让它追踪产生最多可避免费用和税收拖累的重复持仓,每个数据都精确引用到具体的文件和页码——整个过程无需任何文档接触到聊天机器人。
  • 初创公司:询问为何激活率停滞不前,智能体会在本地分析漏斗导出文件,找出从安装到首次完成任务之间新用户流失的位置,然后将主要洞察直接发布到团队的 Slack 频道。

Try Portable Computer today.

立即试用 Portable Computer。

Hermes Agent — developed by Nous Research — is a general-purpose agent used by millions that excels at reliability and self-improvement. Model- and provider-agnostic, Hermes is built to run all day on local systems, making RTX PCs, RTX PRO workstations and DGX Spark a natural fit.

Hermes Agent——由 Nous Research 开发——是一款被数百万用户使用的通用智能体,在可靠性和自我改进方面表现出色。Hermes 不依赖特定模型或提供商,专为在本地系统上全天候运行而设计,使其成为 RTX PC、RTX PRO 工作站和 DGX Spark 的理想选择。

Configuring a local model in Hermes will provide users with one-click setup across RTX and DGX systems on Windows. The agent will automatically detect the NVIDIA GPU, select an appropriate model and configuration, and run it through integrated llama.cpp with NVIDIA inference optimizations already in place, eliminating manual model downloads and tuning. Support for Linux is coming soon.

在 Hermes 中配置本地模型将为用户提供在 Windows 上的 RTX 和 DGX 系统上一键设置的功能。该智能体将自动检测 NVIDIA GPU,选择合适的模型和配置,并通过集成 llama.cpp 运行,其中已内置 NVIDIA 推理优化,从而消除手动下载模型和调整参数的需要。Linux 支持即将推出。

Once it is running, Hermes works the way it does anywhere else. It uses tools, maintains context across tasks, remembers information between sessions and creates reusable skills over time, allowing the agent to become more capable with continued use. Running the model locally on a GPU keeps performance fast while keeping data on the system.

一旦运行起来,Hermes 的工作方式与其他地方一样。它使用工具,跨任务保持上下文,在会话之间记住信息,并随着时间的推移创建可重用的技能,使智能体随着持续使用变得更加强大。在 GPU 上本地运行模型可保持快速性能,同时将数据保留在系统内。

One-click local model setup is available now on Windows, with support coming soon to Linux. Learn more about Hermes Agent.

Windows 上现已提供一键式本地模型设置,Linux 支持即将推出。了解更多关于 Hermes Agent 的信息。

OpenClaw has become one of the defining projects of the open-agent movement — the largest AI project on GitHub, with more than 380K stars and a fast-growing community that’s building tools and skills across research, engineering, project management and everyday productivity.

OpenClaw 已成为开放智能体运动的标志性项目之一——它是 GitHub 上最大的 AI 项目,拥有超过 38 万颗星,并且拥有一个快速增长的社区,正在研究、工程、项目管理以及日常生产力领域构建工具和技能。

NVIDIA, Microsoft and OpenClaw have been working together to make that experience easier to set up on Windows PCs. To reduce onboarding friction, the OpenClaw Windows App simplifies the process of setting up an optimized local model on any RTX GPU with at least 24GB of VRAM.

NVIDIA、Microsoft 和 OpenClaw 一直在合作,使这一体验在 Windows PC 上更容易设置。为了减少入门摩擦,OpenClaw Windows App 简化了在至少配备 24GB VRAM 的任何 RTX GPU 上设置优化本地模型的过程。

Learn more in the OpenClaw blog.

在 OpenClaw 博客中了解更多。

Faster Inference Gives Local Agents a Boost

更快的推理为本地智能体带来提升

Inference performance is critical to keeping local agents responsive. NVIDIA is continuing to collaborate with the open-source llama.cpp and vLLM communities to accelerate agentic workloads across local NVIDIA platforms.

推理性能对于保持本地智能体的响应能力至关重要。NVIDIA 继续与开源 llama.cpp 和 vLLM 社区合作,以加速本地 NVIDIA 平台上的智能体工作负载。

llama.cpp delivers up to 1.9x higher throughput through kernel optimizations on a GeForce RTX 5090, enhanced speculative decoding techniques and faster prefill. ​

llama.cpp 通过 GeForce RTX 5090 上的内核优化、增强的推测解码技术和更快的预填充,实现了高达 1.9 倍的吞吐量提升。

vLLM delivers 1.2x on RTX PRO 6000 Blackwell Workstation Edition and up to 1.4x on two DGX Spark clusters. New XQA attention kernels in FlashInfer and backend optimizations help to accelerate inference across both platforms.

vLLM 在 RTX PRO 6000 Blackwell Workstation Edition 上实现 1.2 倍提升,在两个 DGX Spark 集群上实现高达 1.4 倍提升。FlashInfer 中的新 XQA 注意力内核和后端优化有助于加速这两个平台上的推理。

These gains are available on the llama.cpp and vLLM inferencing backends.

这些增益在 llama.cpp 和 vLLM 推理后端均可用。

Users can also experience these via the LM Studio and Ollama applications.

用户还可以通过 LM Studio 和 Ollama 应用程序体验这些功能。

Tap Idle PCs for More Local AI Compute With NVIDIA PAIR

利用 NVIDIA PAIR 让闲置 PC 发挥更多本地 AI 算力

More than half of U.S. households have two or more PCs, and much of that computing power sits idle throughout the day. NVIDIA Personal AI Router (PAIR) is a free, open source software tool that puts those systems to work together for local AI.

超过一半的美国家庭拥有两台或更多电脑,而这些计算资源在一天中大部分时间都处于闲置状态。NVIDIA Personal AI Router(PAIR)是一款免费开源的软件工具,可将这些系统协同工作以支持本地 AI。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近