跳到主内容
@wquguru
精选88MarkTechPost(RSS)行业动态多源精选 ×5

NVIDIA发布64GB版DGX Spark桌面AI系统

NVIDIA Announces DGX Spark 64GB: A 1-PetaFLOP Grace Blackwell Desktop for Local AI Agents, Fine-Tuning, and Inference

原文
发到 X
推荐理由

英伟达将Grace Blackwell算力下沉至桌面端,为本地Agent和私有数据微调提供了低门槛硬件方案,关注边缘部署与算力硬件的同学值得重点了解。

NVIDIA announced a new 64GB configuration of DGX Spark — from Acer, ASUS, Dell, Gigabyte, HP and MSI — its GB10-powered desktop AI system. It gives developers a way to start with one system for local models and agents, then cluster two 64GB units for 128GB of memory across the cluster and more compute when workloads grow.

NVIDIA 宣布推出由 Acer、ASUS、Dell、Gigabyte、HP 和 MSI 提供的 DGX Spark 全新 64GB 配置——这是一款搭载 GB10 芯片的桌面 AI 系统。它为开发者提供了一种方式:先使用一台系统运行本地模型和智能体,然后在负载增长时,将两台 64GB 单元集群化,从而在集群中获得 128GB 内存和更多算力。

The direct message is: Run open models and always-on agents on your own desk, then cluster DGX Spark system as the work grows – instead of on a metered API.

其核心信息是:在你自己的桌面上运行开源模型和全天候运行的智能体,然后随着工作负载的增长对 DGX Spark 系统进行集群部署——而不是依赖按量计费的 API。

The timing is not accidental. Agents burn tokens continuously through tool calls, retries, long context, and multi-step plans. Token consumption has grown 14x since early 2026. On a cloud API, every one of those tokens is billed. On owned hardware, there is no per-token fee.

这一时机并非偶然。智能体通过工具调用、重试、长上下文和多步计划持续消耗 token。自 2026 年初以来,token 消耗量增长了 14 倍。在云 API 上,每一个 token 都会被计费;而在自有硬件上,则没有按 token 收费的费用。

The 64GB model keeps the same GB10 Grace Blackwell superchip, NVIDIA CUDA accelerated AI software stack, and ConnectX-7 networking as the original — updated with 64GB of unified LPDDR5x instead of 128GB. NVIDIA positions that as enough for today’s most capable 30–35B class open models. NVIDIA continues to offer the 128GB DGX Spark for larger single-box workloads, and clustering 64GB units adds memory and compute together.

这款 64GB 型号保留了与原版相同的 GB10 Grace Blackwell 超级芯片、NVIDIA CUDA 加速 AI 软件栈以及 ConnectX-7 网络功能,不同之处在于将 128GB 统一内存更新为 64GB LPDDR5x。NVIDIA 认为这足以应对当今最强大的 30–35B 参数级别的开源模型。NVIDIA 继续提供 128GB 版本的 DGX Spark 以支持更大的单机负载,而将 64GB 单元集群化则可以同时增加内存和算力。

What is Inside the Box

盒子里有什么

GB10 pairs a Blackwell GPU with 5th-generation Tensor Cores and a 20-core Grace Arm CPU. The GPU delivers up to 1 petaFLOP of FP4 AI compute, with sparsity.

GB10 将配备第五代 Tensor Core 的 Blackwell GPU 与拥有 20 核心的 Grace Arm CPU 配对。该 GPU 可提供高达 1 petaFLOP 的 FP4 AI 算力(支持稀疏计算)。

SpecDGX Spark 64GB
SuperchipNVIDIA GB10 Grace Blackwell
CPU20-core Arm (10× Cortex-X925 + 10× Cortex-A725)
AI computeUp to 1 petaFLOP FP4 (with sparsity)
Memory64GB LPDDR5x, coherent unified
Memory bandwidth273 GB/s
Storage1, 2, or 4TB NVMe M.2, self-encrypting
NetworkingConnectX-7 NIC, 200GbE; Wi-Fi 7, BT 5.3
Display1× HDMI 2.1a
OSNVIDIA DGX OS (Ubuntu-based)
Size / weight150 × 150 × 50.5 mm / 1.2 kg
Max local model sizeUp to 100B parameters
AvailabilityOct 23, 2026 — NVIDIA Marketplace, OEM partners, retail
规格DGX Spark 64GB
超级芯片NVIDIA GB10 Grace Blackwell
CPU20 核 Arm(10× Cortex-X925 + 10× Cortex-A725)
AI 算力最高 1 petaFLOP FP4(支持稀疏计算)
内存64GB LPDDR5x,相干统一内存
内存带宽273 GB/s
存储1、2 或 4TB NVMe M.2,支持自加密
网络ConnectX-7 网卡,200GbE;Wi-Fi 7,蓝牙 5.3
显示1× HDMI 2.1a
操作系统NVIDIA DGX OS(基于 Ubuntu)
尺寸/重量150 × 150 × 50.5 毫米 / 1.2 公斤
最大本地模型大小最高 100B 参数
上市时间2026 年 10 月 23 日 — NVIDIA Marketplace、OEM 合作伙伴、零售渠道

Source: NVIDIA specifications.

来源:NVIDIA 规格说明。

Unified memory is the key design choice: The CPU and GPU share one pool over NVLink-C2C, at 5x the bandwidth of PCIe Gen 5. There is no copying weights between system RAM and VRAM. For agents, that means several models, their KV caches, and tool processes live in one address space.

统一内存是关键的设计选择:CPU 和 GPU 通过 NVLink-C2C 共享同一内存池,其带宽是 PCIe Gen 5 的 5 倍。系统 RAM 和显存之间无需复制权重数据。对于智能体而言,这意味着多个模型、它们的 KV 缓存以及工具进程都存在于同一个地址空间中。

The software is ready on first boot: DGX OS ships with the NVIDIA AI stack, including PyTorch, Jupyter, and Ollama. NVIDIA NemoClaw installs with a single command. It adds privacy and security controls to OpenClaw agents. NVIDIA OpenShellTM, part of the NVIDIA Agent ToolkitTM, adds policy-based guardrails on top. NVIDIA NemotronTM models are optimized for the box.

软件在首次启动时即可就绪:DGX OS 预装了 NVIDIA AI 堆栈,包括 PyTorch、Jupyter 和 Ollama。NVIDIA NemoClaw 可通过一条命令安装,为 OpenClaw 智能体添加隐私与安全控制功能。作为 NVIDIA Agent ToolkitTM 一部分的 NVIDIA OpenShellTM 在此基础上增加了基于策略的安全护栏。NVIDIA NemotronTM 模型已针对该硬件进行优化。

It runs on a standard wall outlet: No server room, no special cooling. That is important for an agent meant to run around the clock.

它可直接接入标准电源插座:无需服务器机房,也无需特殊冷却系统。这对于需要全天候运行的智能体而言至关重要。

Built-in networking for clustering: ConnectX-7 lets two DGX Spark 64GB systems cluster for 128GB of memory and more compute. NVIDIA Sync Cluster Assistant simplifies setup.

内置集群网络支持:ConnectX-7 允许两台 DGX Spark 64GB 系统组成集群,提供 128GB 内存及更多算力。NVIDIA Sync Cluster Assistant 简化了设置流程。

Which Open Models Fit in 64GB

哪些开源模型可适配 64GB 显存

Hardware is half the story. The other half is that 30B-class open models got good enough for agent work.

硬件只是故事的一半。另一半在于,30B 级别的开源模型已足够胜任智能体工作。

ModelDeveloperTypeFootprintRole on a Spark
Muse GlimmerMeta29.6B dense, text + image, Apache 2.0~17GB (quantized)Main agent model
Nemotron 3.5 LightningNVIDIA30B MoE (30B-A3B)NVFP4 checkpointFast executor for long-running agents
Qwen3.8-27BAlibaba Qwen27B dense~13.5GB weights (4-bit)General agent and coding
模型开发者类型占用空间在 Spark 上的角色
Muse GlimmerMeta29.6B 稠密模型,文本+图像,Apache 2.0 协议~17GB(量化后)主智能体模型
Nemotron 3.5 LightningNVIDIA30B MoE(30B-A3B)NVFP4 检查点长运行智能体的快速执行器
Qwen3.8-27BAlibaba Qwen27B 稠密模型~13.5GB 权重(4-bit)通用智能体与代码编写

Footprints are weight-only estimates; KV cache and runtime overhead come on top.

占用空间仅为权重估算值;KV 缓存和运行时开销需额外计算。

Muse Glimmer is the main model. Meta distilled it from Muse Spark, the model family behind the Meta AI assistant. It targets local agents: reliable tool calls, long multi-step tasks, and recovery from failures. Meta reports 51.2 on SWE-Bench Pro and 75.5 on MCP Atlas. Context runs to 131K tokens. It is also available as an NVIDIA NIM.

Muse Glimmer 是主力模型。Meta 从 Muse Spark(Meta AI 助手背后的模型家族)中蒸馏出该模型。它专为本地智能体设计:具备可靠的工具调用能力、支持长多步任务以及故障恢复。Meta 报告其在 SWE-Bench Pro 上得分为 51.2,在 MCP Atlas 上得分为 75.5。上下文窗口可达 131K tokens。该模型也可通过 NVIDIA NIM 获取。

Full BF16 Glimmer needs 55GB+, which leaves almost nothing for context. The ~17GB quantized build is the practical choice on 64GB.

完整的 BF16 精度 Glimmer 需要 55GB+ 显存,几乎不留空间给上下文。在 64GB 设备上,~17GB 的量化版本是实际可行的选择。

Larger models like DeepSeek V4 Flash need more than one box. NVIDIA’s own benchmarks run it on 4 64 GB clustered Sparks.

像 DeepSeek V4 Flash 这样的大型模型需要多台设备。NVIDIA 自身的基准测试显示,需在 4 台 64GB 集群配置的 Spark 上运行该模型。

5 Things You Can Build on One Box

单台设备可构建的五种应用

1. An always-on personal agent

1. 始终在线的个人智能体

Install Hermes and point it at Muse Glimmer or Nemotron 3.5 Lightning. Give it tools: your GitHub repos, a test runner, an RSS feed of arXiv categories. Let it run overnight.

安装 Hermes 并指向 Muse Glimmer 或 Nemotron 3.5 Lightning。为其配置工具:你的 GitHub 仓库、测试运行器、arXiv 分类的 RSS 订阅源。让它整夜运行。

By morning it has triaged new issues, reproduced a failing test, and drafted a pull request for review. It has also summarized the 30 papers you would never have opened. Your private notes, code, and email never leave the machine.

到第二天早上,它已完成新问题的分类、复现了失败的测试,并起草了待审查的拉取请求。它还总结了你永远不会打开阅读的 30 篇论文。你的私人笔记、代码和邮件均不会离开本机。

2. Fine-tune a coding model on your own repo

2. 在你的私有仓库上微调代码模型

QLoRA on a 70B model fits in 64GB. Train it on your codebase, internal docs, and past PR reviews. Then serve it locally as a coding assistant that knows your conventions.

在 70B 模型上使用 QLoRA 可适配于 64GB 内存。在你的代码库、内部文档和过往 PR 审查记录上进行训练。然后在本地部署为编码助手,使其了解你的开发规范。

NVIDIA measured ~18,400 tokens/s on a single node for nanochat distributed fine-tuning.

NVIDIA 在单节点上对 nanochat 分布式微调测得约 18,400 tokens/s 的速度。

3. A day-1 model evaluation bench

3. 首日模型评估基准

A new open-weight model drops on Hugging Face. Pull it through Ollama or vLLM the same day. Run your own question set against it.

一个新的开源权重模型在 Hugging Face 发布。当天即可通过 Ollama 或 vLLM 拉取。使用你自己的问题集对其进行测试。

Score accuracy, then measure time to first token (TTFT) and tokens per second. Compare it against your current model. There is no API bill for re-running the eval 50 times.

先评估准确率,然后测量首 token 时间(TTFT)和每秒 token 数。将其与你当前的模型进行比较。重新运行 50 次评估无需支付 API 费用。

4. A multi-model agent team

4. 多模型智能体团队

Unified memory lets several models share one pool. Run Bonsai 2 as a router, and Glimmer as the main reasoning agent. That is roughly 23GB of weights.

统一内存让多个模型共享一个池。运行 Bonsai 2 作为路由器,Glimmer 作为主要推理智能体。这大约占用 23GB 的权重空间。

The rest covers KV cache, the OS, and tool processes. When a problem exceeds local capacity, the agent can send a sanitized question to a larger cloud model.

其余部分用于 KV 缓存、操作系统和工具进程。当问题超出本地处理能力时,智能体可将脱敏后的问题发送给更大的云端模型。

5. Edge and robotics prototyping

5. 边缘与机器人原型设计

Fine-tune a vision transformer for a specific task, like spotting anomalies on a factory floor camera. Validate it locally with the same CUDA stack. Then deploy it to an NVIDIA Jetson device at the edge.

针对特定任务微调视觉 Transformer,例如在工厂车间摄像头中检测异常。使用相同的 CUDA 栈在本地进行验证。然后将其部署到边缘的 NVIDIA Jetson 设备上。

When One Box is Not Enough: Clustering with NVIDIA Sync

当 One Box 不够用时:使用 NVIDIA Sync 进行集群化

Every DGX Spark ships with ConnectX-7 at 200GbE. Clustering is a native feature, not an add-on.

每台 DGX Spark 均配备 200GbE ConnectX-7。集群化是原生功能,而非附加组件。

SetupPooled memoryAI compute (FP4)Connection
1 Spark (64GB)64GBUp to 1 PFLOP—
2 Sparks (64GB each)128GBUp to 2 PFLOPSDirect QSFP cable, no switch
2 Sparks (128GB each)256GBUp to 2 PFLOPSDirect QSFP cable, no switch
3 Sparks (128GB each)384GBUp to 3 PFLOPSQSFP ring, no switch
4 Sparks (128GB each)512GBUp to 4 PFLOPS200GbE switch (QSFP56-DD, RoCE v2)
设置池化内存AI 计算 (FP4)连接
1 个 Spark (64GB)64GB高达 1 PFLOP—
2 个 Spark (各 64GB)128GB高达 2 PFLOPS直连 QSFP 线缆,无交换机
2 个 Spark (各 128GB)256GB高达 2 PFLOPS直连 QSFP 线缆,无交换机
3 个 Spark (各 128GB)384GB高达 3 PFLOPSQSFP 环网,无交换机
4 个 Spark (各 128GB)512GB高达 4 PFLOPS200GbE 交换机 (QSFP56-DD, RoCE v2)

Source: NVIDIA. 3-node compute derived from per-unit specs.

来源:NVIDIA。3 节点计算基于单机规格推导得出。

NVIDIA states 2 clustered 64GB units deliver up to 1.7x the performance of one 128GB DGX Spark. The reason is doubled AI compute and bandwidth: 2 units provide up to 546 GB/s combined, double a single box.

NVIDIA 指出,2 个集群化的 64GB 单元可提供比单个 128GB DGX Spark 高 1.7 倍的性能。原因在于 AI 计算能力和带宽翻倍:2 个单元提供高达 546 GB/s 的总带宽,是单台设备的两倍。

NVIDIA Sync is the management layer. The Windows and macOS app discovers Sparks on your network and manages SSH access. Its Cluster Assistant configures ConnectX-7 networking for up to 4 systems. Nodes can also connect across locations over a Tailscale mesh, with no cloud in the data path.

NVIDIA Sync 是管理层。Windows 和 macOS 应用可发现网络中的 Spark 设备并管理 SSH 访问。其 Cluster Assistant 可为多达 4 套系统配置 ConnectX-7 网络。节点还可通过 Tailscale 网格跨地点连接,数据路径中无需经过云端。

The pattern is quite clear. Clustering roughly halves time to first token per doubling and scales fine-tuning near-linearly. Decode improves more modestly, about 1.4x at 4 nodes. For agents that read long inputs, the TTFT gain is the one that matters.

模式非常清晰。聚类(Clustering)使首次令牌生成时间(TTFT)随节点数翻倍而大致减半,且微调扩展性接近线性。解码性能提升较为温和,在4个节点时约为1.4倍。对于读取长输入的智能体而言,TTFT的增益才是关键所在。

Step-by-step multi-node guides, including vLLM on stacked Sparks, are at build.nvidia.com/spark.

逐步式多节点指南,包括在堆叠Spark上的vLLM部署,请访问 build.nvidia.com/spark。

What It is, and What It is Not

它是什么,以及不是什么

  • Not a chat server for 100 users: 273 GB/s of bandwidth limits concurrent decode throughput.
  • Built for long-input, short-output work: Reading a repo, a log dump, or a paper stack, then writing a short result.
  • 64GB caps one box at 100B parameters: For more, cluster 2 units or choose the 128GB configuration.
  • 并非面向100名用户的聊天服务器:273 GB/s的带宽限制了并发解码吞吐量。
  • 专为长输入、短输出工作负载打造:例如阅读代码仓库、日志转储或论文堆栈,然后生成简短结果。
  • 64GB内存将单台设备限制在100B参数以内:如需更大规模,请集群部署2个单元或选择128GB配置。

Key Takeaways

关键要点

  • DGX Spark 64GB keeps GB10, 1 PFLOP FP4, and the full NVIDIA AI stack.
  • 64GB runs today’s 30–35B class open models, like Qwen 3.8 27B and Nemotron 3.5 Lightning.
  • Best uses: always-on agents, QLoRA fine-tuning, and day-1 model evals — no per-token fees.
  • DGX Spark 64GB保留了GB10架构、1 PFLOP FP4算力以及完整的NVIDIA AI软件栈。
  • 64GB内存可运行当今30–35B级别的开源模型,如Qwen 3.8 27B和Nemotron 3.5 Lightning。
  • 最佳应用场景:始终在线的智能体、QLoRA微调以及首日模型评估——无需按令牌付费。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →