跳到主内容
@wquguru
精选75Latent Space(RSS)行业动态多源精选 ×2

内存价格一年暴涨500%,2027年DRAM产能被抢空

[AINews] Memory prices up 500% in 12 months

原文
发到 X

Even as Sama follows through on the Great Pacing, and Etched becomes a double unicorn and Cerebras announced CS4 running 10T models at 1000 tok/s, the memory shortage has continued unabated since we did our SemiAnalysis pod in Feb.

即便Sama持续推进“伟大节奏”,Etched成为双独角兽,Cerebras宣布CS4以1000 tok/s运行10T模型,自我们2月发布SemiAnalysis播客以来,内存短缺问题仍未缓解。

Per Tom’s Hardware:

据Tom's Hardware报道:

We’re officially in dire straits. There’s almost no way, if you’re reading this site, that you aren’t aware that memory prices have become entirely divorced from reality. Some are calling it the RAMpocalypse; I prefer “RAMageddon.”

我们正式陷入困境。如果你正在阅读本网站,你几乎不可能不知道内存价格已完全脱离现实。有人称之为“内存末日”;我更倾向于“内存浩劫”。

That’s right: 128GB DDR5 kits are fully ten times more expensive than the lowest price we’ve ever seen.

没错:128GB DDR5套件的价格已是我们所见最低价的整整十倍。

In fact, the situation is so severe that hyperscale buyers have reportedly already locked in almost all of the global DRAM production capacity for 2027, handing over advance deposits to guarantee their supply of precious DRAM, which is now among the highest-value commodities in the world by weight; mainstream DRAM chips are worth over half as much per kilogram as solid gold.

事实上,情况如此严峻,据报道,超大规模买家已锁定2027年全球DRAM产能的绝大部分,并预付定金以确保其宝贵的DRAM供应。按重量计算,DRAM现已成为全球价值最高的商品之一;主流DRAM芯片每公斤价值超过纯金的一半。

Put another way, the famous Moore’s Law driving all hardware unit prices down has been reversed for memory:

换句话说,推动所有硬件单价下降的著名摩尔定律在内存领域已被逆转:

AI News for 8/17/2026-8/18/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!

2026年8月17日至8月18日的AI新闻。我们检查了12个子版块、544个Twitter账号,没有其他Discord。AINews网站可搜索所有过往期次。提醒一下,AINews现已成为Latent Space的一部分。你可以选择接收或退订邮件频率!

AI Twitter Recap

AI Twitter摘要

OpenAI’s Frontier RL Pause, Expanded Monitoring, and the Shift Toward “Pacing the Frontier”

OpenAI的前沿强化学习暂停、扩展监控,以及向“前沿节奏”的转变

  • OpenAI slowed frontier training to harden security and alignment controls: The day’s biggest systems/safety development was OpenAI saying it paused some frontier RL training for two weeks and is still holding its largest planned frontier RL run while it strengthens monitoring, isolation, and red-teaming. Sam Altman framed this as a case where capabilities were outpacing safety/alignment readiness, while Greg Brockman emphasized that confidence in safety will increasingly set the pace of frontier scaling. OpenAI also clarified the slowdown mainly affects farther-out releases, not models already near ship.
  • Concrete controls matter more than broad messaging: OpenAI shared more implementation detail than usual, including stronger workload/network isolation, continuous security testing, and multistage monitoring. Secondary commentary highlighted interesting operational details: monitoring may add roughly 20% overhead, sampled-token monitoring can page safety/security/research teams within ~30 minutes, and tool-using inference for higher-risk systems may ship with active monitors attached, per @eliebakouch. Whatever one thinks of the policy framing, this is notable as a public admission that training/eval infra and inference-time monitors are now bottlenecks on frontier progress, not just raw compute.
  • OpenAI放缓前沿训练以加强安全与对齐控制:当天最大的系统/安全进展是OpenAI表示暂停部分前沿强化学习训练两周,并在加强监控、隔离和红队测试的同时,仍保持其最大规模的前沿强化学习计划。Sam Altman将此描述为能力超越安全/对齐准备的情况,而Greg Brockman强调,对安全的信心将日益决定前沿扩展的节奏。OpenAI还澄清,放缓主要影响更远期发布的版本,而非接近发布的模型。
  • 具体控制措施比宽泛的信息传递更重要:OpenAI 分享了比以往更多的实施细节,包括更强的工作负载/网络隔离、持续的安全测试和多阶段监控。次要评论强调了有趣的运营细节:监控可能增加约20%的开销,采样令牌监控可在约30分钟内通知安全/安保/研究团队,而针对高风险系统的工具使用推理可能会附带主动监控器发布,据@eliebakouch所述。无论人们对政策框架持何种看法,这都值得注意,因为这是公开承认训练/评估基础设施和推理时监控现在成为前沿进展的瓶颈,而不仅仅是原始计算力。

Open Models: Qwen3.8-27B Momentum, GLM-5.3’s Post-Training Gains, and the Small-Model Debate

开放模型:Qwen3.8-27B Momentum、GLM-5.3 的后训练收益,以及小模型之争

  • Qwen3.8-27B became the focal point of the local/open model conversation: Several posts cast Qwen3.8-27B as a new “locally runnable frontier-ish” moment, with @kimmonismus calling it a “DeepSeek moment” and Alibaba Qwen celebrating it reaching #1 local model in Cline in four days. Benchmarks cited in the thread include #7 on Artificial Analysis’ Agentic Index at 27B, #6 among open-weight models on Vals Index v2 and #1 on Harvey’s legal benchmark among open weights, and Cline’s own ranking as its new top local model. The pushback was equally strong: @scaling01 argued benchmark wins are overstated versus Opus 4.5 in real coding use, underscoring the growing divide between bench success, cost efficiency, and qualitative reliability on long tasks.
  • Safety implications of capable local models are getting harder to dismiss: A high-engagement post from @kimmonismus noted a “refusal-removed” MLX build of Qwen3.8-27B running locally on Apple Silicon in 2/4/6/8-bit variants, claiming preserved vision, reasoning, tool use, and 262K context with near-zero refusals. Independent of the rhetoric, this is the clearest thread in the set pointing to a real shift: useful, locally deployable, partially uncensored models are no longer hypothetical.
  • GLM-5.3 looks like a post-training/infrastructure story, not a base-model story: Z.ai launched GLM-5.3 via API for coding, defensive cyber, and long-horizon agents, at the same price as GLM-5.2. Artificial Analysis reported it ties Kimi K3 at 60 on its Intelligence Index, with a 246-point jump on GDPval-AA v2 to 1770 Elo, while keeping the same 753B total / 40B active MoE footprint, 1M context, and MIT license once weights land. The most technically interesting interpretation came from a long Zhihu summary relayed by @ZhihuFrontier: GLM-5.3’s gains appear driven by stronger post-training, especially asynchronous RL (SAO), executable sandbox training, and on-policy distillation to prevent catastrophic forgetting. If true, this is a meaningful data point for the idea that agentic capability scaling is shifting from parameter count toward RL systems + environment quality.
  • Qwen3.8-27B 成为本地/开放模型讨论的焦点:多篇帖子将 Qwen3.8-27B 视为新的“本地可运行前沿级”时刻,@kimmonismus 称其为“DeepSeek 时刻”,阿里巴巴 Qwen 则庆祝其在四天内成为 Cline 中排名第一的本地模型。帖子中引用的基准包括在 Artificial Analysis 的 Agentic Index 上以 27B 规模排名第7,在 Vals Index v2 上位列开放权重模型第6,在 Harvey 的法律基准上位列开放权重模型第1,以及 Cline 自身排名将其列为新的顶级本地模型。反对声音同样强烈:@scaling01 认为,在真实编码使用中,基准胜利相对于 Opus 4.5 被夸大了,这凸显了基准成功、成本效益和长任务定性可靠性之间日益扩大的差距。
  • 强大本地模型的安全影响越来越难以忽视:@kimmonismus 的一篇高互动帖子指出,Qwen3.8-27B 的“去拒绝”MLX 构建可在 Apple Silicon 上以 2/4/6/8 位变体本地运行,声称保留了视觉、推理、工具使用和 262K 上下文,且几乎零拒绝。撇开修辞不谈,这是该组中最清晰的线索,指向一个真实的转变:有用的、可本地部署的、部分未经审查的模型不再是假设性的。
  • GLM-5.3 看起来更像是一个后训练/基础设施的故事,而非基础模型的故事:Z.ai 通过 API 发布了 GLM-5.3,面向编程、防御性网络和长周期智能体,价格与 GLM-5.2 相同。Artificial Analysis 报告称,其智能指数与 Kimi K3 并列 60 分,在 GDPval-AA v2 上跃升 246 分至 1770 Elo,同时保持相同的 753B 总参数量 / 40B 激活 MoE 架构、1M 上下文,权重发布后采用 MIT 许可证。最有趣的技术解读来自 @ZhihuFrontier 转发的一篇长篇知乎总结:GLM-5.3 的提升似乎主要源于更强的后训练,尤其是异步强化学习(SAO)、可执行沙箱训练,以及在线策略蒸馏以防止灾难性遗忘。如果属实,这为“智能体能力扩展正从参数量转向强化学习系统 + 环境质量”这一观点提供了有意义的数据点。

Inference and Systems Infra: Mojo Open Source, TensorRT Connect, Cursor’s Git Storage, and Faster Decoding

推理与系统基础设施:Mojo 开源、TensorRT Connect、Cursor 的 Git 存储,以及更快的解码

  • Mojo is now open source under Apache 2.0: Modular’s announcement drew broad attention, with the company formally open-sourcing Mojo and also positioning its broader platform as a portability layer across accelerators, including Qualcomm datacenter AI accelerators. For infra engineers, the significance is less “new language hype” than toolchain openness plus hardware abstraction arriving together.
  • NVIDIA compressed model-to-TensorRT deployment to “two commands”: NVIDIA launched TensorRT Model Connect in public preview, promising direct conversion from supported Hugging Face models to end-to-end TensorRT inference without intermediate ONNX export, with output deployable via native C++ APIs. The post also claims the project itself was largely built with Codex agents under human review, which is noteworthy less as marketing than as another signal that infra/tooling teams are now willing to say agent assistance touched implementations, tuning, tests, integrations, and docs.
  • Cursor published a strong infra retrospective on Git hosting at scale: The standout systems post by engagement was Cursor’s writeup on designing Git storage “as if it were a database”. This is adjacent to AI rather than model-specific, but highly relevant for anyone building coding-agent backends: as agents amplify repo churn, background automation, and branch/session proliferation, Git hosting becomes a core AI infra dependency rather than a generic devops primitive.
  • Fast decoding and accelerator claims kept escalating: On-device inference got a notable boost with DFlash 2 claiming Qwen3.8-27B at 70 tok/s on an M5 Max, up to 4.6× autoregressive decoding “with the same output.” On the datacenter side, Cerebras announced CS-4, with follow-on claims around 10T models at 1000 tok/s, ~1300 tok/s for GPT-5.6 Sol, and up to 10× higher throughput per MW. Even allowing for vendor framing, the throughline is clear: inference speed is becoming product UX, economics, and national-competitiveness policy all at once.
  • Mojo 现已在 Apache 2.0 下开源:Modular 的公告引起了广泛关注,公司正式开源 Mojo,并将其更广泛的平台定位为跨加速器的可移植层,包括高通数据中心 AI 加速器。对于基础设施工程师而言,其意义更多在于工具链开放性与硬件抽象同时到来,而非“新语言炒作”。
  • NVIDIA 将模型到 TensorRT 的部署压缩为“两条命令”:NVIDIA 推出了 TensorRT Model Connect 公开预览版,承诺可直接从受支持的 Hugging Face 模型转换为端到端的 TensorRT 推理,无需中间 ONNX 导出,输出可通过原生 C++ API 部署。该文章还声称该项目本身主要是在人工审查下由 Codex 智能体构建的,这与其说是营销,不如说是另一个信号,表明基础设施/工具团队现在愿意承认智能体辅助涉及了实现、调优、测试、集成和文档。
  • Cursor 发布了一篇关于大规模 Git 托管的深度基础设施回顾:按参与度衡量,最突出的系统文章是 Cursor 关于“像设计数据库一样”设计 Git 存储的论述。这与 AI 相邻而非模型特定,但对任何构建编程智能体后端的人来说都高度相关:随着智能体放大仓库变更、后台自动化和分支/会话的激增,Git 托管成为核心 AI 基础设施依赖,而非通用的开发运维原语。
  • 快速解码和加速器的声明不断升级:设备端推理因 DFlash 2 声称在 M5 Max 上以 70 tok/s 运行 Qwen3.8-27B 而显著提升,自回归解码速度“在相同输出下”最高提升 4.6 倍。在数据中心方面,Cerebras 发布了 CS-4,随后声称 10T 模型可达 1000 tok/s,GPT-5.6 Sol 约 1300 tok/s,每兆瓦吞吐量最高提升 10 倍。即使考虑到供应商的表述,主线依然清晰:推理速度正同时成为产品用户体验、经济性和国家竞争力政策的关键。

Agent Harnesses, Evals, and Production Feedback Loops

智能体工具链、评估与生产反馈循环

  • Miles v0.1 is a serious new OSS RL stack for LLMs and multimodal models: @radixark announced Miles, an open-source RL framework built over 9 months, with 72 contributors, 1,326 commits, and 85 GPU E2E CI tests, reportedly battle-tested on models including Kimi K3, DeepSeek V4, Qwen 3.8, GLM 5.2, Inkling, and MiniMax H3. The pitch is practical: getting RL runs started is easy, but debugging correctness, utilization, and scale is the real bottleneck. This fits the broader theme of the day: the frontier is shifting from “who has PPO/GRPO” to who has robust rollouts, CI, observability, and environment plumbing.
  • Miles v0.1 是一个针对 LLM 和多模态模型的全新严肃开源 RL 栈:@radixark 宣布了 Miles,一个历时 9 个月构建的开源 RL 框架,拥有 72 位贡献者、1,326 次提交和 85 个 GPU 端到端 CI 测试,据称已在包括 Kimi K3、DeepSeek V4、Qwen 3.8、GLM 5.2、Inkling 和 MiniMax H3 等模型上经过实战检验。其定位很务实:启动 RL 运行很容易,但调试正确性、利用率和规模才是真正的瓶颈。这符合当天更广泛的主题:前沿正从“谁拥有 PPO/GRPO”转向谁拥有稳健的 rollout、CI、可观测性和环境管道。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近