跳到主内容
@wquguru
精选90r/LocalLLaMA(Reddit)技巧与观点

用 iPhone 扩展 MacBook 显存加速 Qwen3.8-27B 推理

I made my iPhone a second GPU for my 24 GB MacBook: Qwen 3.8 27B prefills 29–44% faster & my holds part of the CTX window.

原文
发到 X
推荐理由

本地部署同学必看,这套跨设备协同方案巧妙解决了小显存跑大模型的痛点,预填充提速显著且代码开源,值得直接拿来压测你的 Agent 链路。

DISCLAIMER THE PREFILLING TPS SHOWN ON THE PHONE IS COMPUTED ONLY FOR THE LAYERS IT HOLDS. ALREADY FIXING IT TO SHOW END-TO-END PREFILL RATE. NUMBERS BELOW ARE ACCURATE FOR E2E PREFILL RATE.

免责声明:手机上显示的预填充 TPS 仅基于其承载的层计算。目前正在修复以显示端到端预填充速率。以下数字为端到端预填充速率的准确值。

Every file or tool result my agent reads on a 24 GB M4 Pro MacBook is a wait, and 64k of 8-bit context is all that fits next to Qwen 3.8 27B (IQ4_XS), even with the wired limit raised to 20480. An iPhone 17 Pro Max was sitting in my pocket, so I figured what can I do to make use of this extra silicon.

我的代理在配备 24 GB 内存的 M4 Pro MacBook 上读取的任何文件或工具结果都需要等待,而 64k 的 8 位上下文是 Qwen 3.8 27B (IQ4_XS) 旁边能容纳的全部空间,即使将有线限制提高到 20480 也是如此。我的口袋里放着一部 iPhone 17 Pro Max,所以我想我能做些什么来利用这块额外的硅片。

Turns out a 10 Gb/s USB-C cable & some software is all you need. The Mac runs layers 1–40 of each 256-token batch and streams the activations to the phone. The phone runs layers 41–64 on its GPU while the Mac starts the next batch. The A19 Pro's GPU has matrix units (Metal 4 tensor ops), and they make the phone's half 2.4x faster than the same phone without them.

事实证明,一根 10 Gb/s 的 USB-C 线缆和一些软件就足够了。Mac 运行每个 256-token 批次的第 1–40 层,并将激活值流式传输到手机。手机在其 GPU 上运行第 41–64 层,同时 Mac 启动下一个批次。A19 Pro 的 GPU 拥有矩阵单元(Metal 4 tensor ops),这使得搭载该单元的手机的半精度性能比没有这些单元的同一款手机快 2.4 倍。

Same build, phone off vs. on, prefilling a 2,000-token file into a saved agent session:

相同构建,手机关闭与开启状态下,将 2,000-token 的文件预填充到已保存的代理会话中:

  • 8k context: Mac alone 132 tok/s → Mac + iPhone 177 tok/s (+35%) (measured two days earlier, same bench)
  • 16k context: Mac alone 109 tok/s → Mac + iPhone 157 tok/s (+44%)
  • 32k context: Mac alone 101 tok/s → Mac + iPhone 130 tok/s (+29%)
  • 48k context: Mac alone 87 tok/s → Mac + iPhone 113 tok/s (+30%)
  • 8k 上下文:仅 Mac 132 tok/s → Mac + iPhone 177 tok/s (+35%)(两天前测量,基准测试相同)
  • 16k 上下文:仅 Mac 109 tok/s → Mac + iPhone 157 tok/s (+44%)
  • 32k 上下文:仅 Mac 101 tok/s → Mac + iPhone 130 tok/s (+29%)
  • 48k 上下文:仅 Mac 87 tok/s → Mac + iPhone 113 tok/s (+30%)

A fresh 27k-token agent session, cold: 245 s on stock llama.cpp, 228 s on my fork with the Mac alone, and 168 s with the phone.

一个新的 27k-token 代理会话,冷启动状态:使用原版 llama.cpp 耗时 245 秒,仅使用 Mac 时在我修改的版本中耗时 228 秒,加上手机后耗时 168 秒。

Past 64k the phone switches jobs. The oldest KV pages move to the phone and the Mac runs all 64 layers. For every attention layer, the phone computes attention over the old keys on its GPU, and the Mac merges that with its own part. While writing, the phone's Neural Engine takes part of that work too: each 16k-key page of old context is compiled into a Neural Engine model with the keys as its weights. At 140k that took writing from 279 to 176 ms per token compared with the phone's GPU alone.

超过 64k 后,手机切换任务。最旧的 KV 页面被移至手机,Mac 运行全部 64 层。对于每个注意力层,手机在其 GPU 上对旧键计算注意力,Mac 将其与自身部分合并。在写入过程中,手机的神经引擎也参与了部分工作:每个包含 16k 个键的旧上下文页面都被编译为一个神经引擎模型,其中键作为其权重。在 140k 时,这使得每 token 的写入时间从 279 ms 减少到 176 ms,相比仅使用手机 GPU 的情况。

The server allocates 196k–229k of 8-bit context based on the phone's free memory; that's up to ~5.7 GB of KV cache living on the phone instead of the Mac, so the Mac's memory use stops growing at 64k. I've tested a growing session to 128k at 8-bit, with 3/3 planted facts recalled. Separately, at 140k in 4-bit, the run passed the gate with greedy output matching the Mac-only run for 32 generated tokens.

服务器根据手机的可用内存分配 196k–229k 的 8 位上下文;这意味着最多约 5.7 GB 的 KV 缓存存储在手机上而非 Mac 上,因此 Mac 的内存使用量在达到 64k 后停止增长。我已测试过扩展至 128k 的 8 位会话,植入了 3/3 的事实均被回忆起来。另外,在 140k 的 4 位情况下,运行通过了门槛,贪婪输出与前 32 个生成 token 与仅使用 Mac 的运行结果匹配。

What it doesn't do: speed up writing below 64k. That's the Mac's job. My fork's kernels (SME2 on the M4 CPU and Metal fusions) plus DFlash2 speculative decoding take it from 11.3 tok/s on stock llama.cpp to 25 tok/s at about 30k context with medium thinking, phone or not. SME2 also adds up to 29% to prefill on the Mac alone. Past 64k the phone does share the writing (attention over the old keys), and without it the Mac would have to drop to 4-bit context to reach 128k. In real use I have seen upwards of 30 TPS at lower context.

它做不到的事:在低于 64k 时加速写入。那是 Mac 的工作。我的分支的内核(M4 CPU 上的 SME2 和 Metal fusions)加上 DFlash2 推测解码,将其从 stock llama.cpp 的 11.3 tok/s 提升到约 30k 上下文、中等思考时的 25 tok/s,无论是否使用手机。SME2 还单独在 Mac 上将预填充(prefill)速度提升了高达 29%。超过 64k 后,手机会分担写入任务(对旧键进行注意力计算),如果没有它,Mac 必须将上下文降至 4-bit 才能到达 128k。在实际使用中,我在较低上下文下见过超过 30 TPS 的速度。

The phone joins prefills over about 512 tokens. In one real session, that was 7 of 36 requests, but about 83% of the tokens read. Past 64k it holds the context and does the old-key attention, but it stops running layers 41–64 there for now; doing both is next. One request at a time.

手机参与大约 512 个 token 以上的预填充。在一次实际会话中,这是 36 次请求中的 7 次,但占读取 token 的约 83%。超过 64k 后,它保持上下文并执行旧键注意力计算,但目前它停止在那里运行第 41–64 层;同时运行两者是下一步计划。一次只处理一个请求。

I'm curious what this setup could do with newer model architectures. DeepSeek V4.1-Flash reports 890 bytes per token for its global KV cache and adds n-gram embedding tables (Engram). Qwen3.8-Flash-Next, the Qwen 4 architecture preview, has an n-gram lookup table too. Those aren't features of the 27B model I tested, and I haven't benchmarked either architecture here. The real gold is within the newer phones and models working together. With the A20 Pro in the iPhone 18 Pro Max, I bet there is a lot more for me to push.

我很好奇这种设置配合较新的模型架构能做什么。DeepSeek V4.1-Flash 报告其全局 KV 缓存每 token 为 890 字节,并添加了 n-gram 嵌入表(Engram)。Qwen3.8-Flash-Next(Qwen 4 架构预览版)也有一个 n-gram 查找表。这些不是我测试的 27B 模型的特性,我也没有在此基准测试这两种架构。真正的黄金在于新手机和新模型的协同工作。有了 iPhone 18 Pro Max 中的 A20 Pro,我打赌还有更多我可以挖掘的潜力。

Code, setup and bench scripts: https://github.com/StayLameBro/backburner

代码、设置和基准脚本:https://github.com/StayLameBro/backburner

Still a lot of work to do but I built this with Opus 5.5. Happy to answer anything.

仍有很多工作要做,但我用 Opus 5.5 构建了它。乐意回答任何问题。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件