跳到主内容
@wquguru
精选85Hacker News Best(web_list)行业动态多源精选 ×3

AMD收购AI芯片初创Taalas,将模型蚀刻进硅片提升推理性能

AMD收购AI芯片初创公司Taalas,将模型蚀刻进硅片提升推理性能

原文
发到 X
推荐理由

AMD收购Taalas是AI芯片产业链的重大并购,直接影响推理性能与成本,做AI基础设施和模型部署的同学值得关注,建议跟踪后续产品落地。

AI and ML

AI 与机器学习

AMD acquires AI chip startup Taalas to boost inference performance by etching models into silicon

AMD 收购 AI 芯片初创公司 Taalas,通过将模型蚀刻到硅片中提升推理性能

Early tech demos show model-specific integrated circuits churning out up to 17,000 tokens a second

早期技术演示显示,模型专用集成电路每秒可处理高达 17,000 个令牌

Tobias Mann Tobias Mann SYSTEMS EDITOR

Tobias Mann Tobias Mann 系统编辑

Published thu 6 Aug 2026 // 21:05 UTC

发布于 2026 年 8 月 6 日(星期四)// 21:05 UTC

In AMD’s latest bid to upset Nvidia's dominance in AI hardware, the House of Zen has acquired AI chip company Taalas, which bakes model weights directly into silicon in a process that promises to boost inference performance by an order of magnitude or more.

在 AMD 最新一次挑战 Nvidia 在 AI 硬件领域主导地位的尝试中,这家“禅之家”收购了 AI 芯片公司 Taalas,该公司将模型权重直接蚀刻到硅片中,这一过程有望将推理性能提升一个数量级或更多。

The deal, announced at market close on Thursday, appears to be framed in much the same context as Nvidia’s $20 billion licensing deal with Groq last December: make high-performance “premium” inference services prized for AI agents, like code assistants, faster and cheaper to run. AMD didn’t disclose the terms of the deal, but from what we understand, this is an actual acquisition rather than an acquihire.

该交易于周四收盘时宣布,其背景似乎与 Nvidia 去年 12 月与 Groq 达成的 200 亿美元授权协议类似:让面向 AI 代理(如代码助手)的高性能“高级”推理服务运行得更快、成本更低。AMD 未披露交易条款,但据我们所知,这是一次实际收购,而非人才收购。

Founded in 2023 and based in Toronto, Taalas’ approach to inference is radically different from conventional GPUs or the dataflow architectures that underpin Groq LPUs or Cerebras' waferscale accelerators.

Taalas 成立于 2023 年,总部位于多伦多,其推理方法与传统 GPU 或支撑 Groq LPU 和 Cerebras 晶圆级加速器的数据流架构截然不同。

REG AD

REG AD

A model-specific integrated circuit

模型专用集成电路

REG AD

REG AD

The startup’s chips don’t rely on HBM to store the model weights but rather etch them directly into the silicon. In a sense, Taalas’ chips are really model-specific integrated circuits or MSICs.

该初创公司的芯片不依赖 HBM 存储模型权重,而是将权重直接蚀刻到硅片中。从某种意义上说,Taalas 的芯片实际上是模型专用集成电路(MSIC)。

Perhaps more importantly, Taalas’ tech isn’t just conceptual. In February, the startup revealed its first test chip fabbed on TSMC’s 6nm process tech, which it called the HC1. Initial benchmarks saw the chip serve Meta’s Llama 3.1 8B at a blistering 16,960 tokens a second — when announced last February, that was 48x faster than Nvidia's GPUs and 8.5x faster than Cerebras' accelerators.

也许更重要的是,Taalas 的技术并非只是概念。今年 2 月,该初创公司发布了其首款采用台积电 6nm 工艺制造、名为 HC1 的测试芯片。初步基准测试显示,该芯片以惊人的每秒 16,960 个令牌的速度运行 Meta 的 Llama 3.1 8B 模型——在去年 2 月发布时,这比 Nvidia 的 GPU 快 48 倍,比 Cerebras 的加速器快 8.5 倍。

While Llama 3.1 is ancient by today’s standards, having made its debut all the way back in mid 2024, the reticle-sized chip was really intended to prove the concept.

虽然以今天的标准来看,Llama 3.1 已经过时(它早在 2024 年年中就首次亮相),但这款掩模版尺寸的芯片实际上是为了验证这一概念。

Taalas has been incredibly secretive about how its chips actually work, but we know its processors are comprised of two main regions: the mask-ROM recall fabric where model weights are etched, and the SRAM recall fabric where KV caches and fine-tuning adapters are stored.

Taalas 对其芯片的实际工作原理一直守口如瓶,但我们知道其处理器由两个主要区域组成:掩模 ROM 召回结构(模型权重蚀刻于此)和 SRAM 召回结构(KV 缓存和微调适配器存储于此)。

For its second-gen HC2 chip due out this summer, Taalas aims to boost parameter count to 20 billion parameters. That might not sound like much, but just like with GPUs for larger models, weights are simply distributed across multiple accelerators using pipeline parallelism.

对于计划于今年夏天推出的第二代 HC2 芯片,Taalas 的目标是将参数数量提升至 200 亿。这听起来可能不多,但就像 GPU 处理更大模型一样,权重可以通过流水线并行简单地分布在多个加速器上。

At 20 billion parameters per chip, you’d need just 50 accelerators to support a trillion-parameter model, and AMD just so happens to have a rack-scale compute platform and in-house system design team that can comfortably accommodate that.

每颗芯片拥有200亿参数,支持一个万亿参数模型仅需50个加速器,而AMD恰好拥有机架级计算平台和内部系统设计团队,可以轻松满足这一需求。

That’s quite a bit more space and power efficient than Nvidia’s recently unveiled LPX systems, which would need a few dozen GPUs and at least 2,000 Groq LPUs to serve the same model.

这比英伟达最近推出的LPX系统在空间和功耗效率上要高得多,后者需要几十个GPU和至少2000个Groq LPU才能服务相同的模型。

From what we understand, AMD intends to pair its Instinct-based Helios racks with chips based on Taalas’ tech, which implies a disaggregated architecture where compute-heavy prompt processing is done on GPUs while token generation is offloaded to Taalas-based accelerators.

据我们所知,AMD计划将其基于Instinct的Helios机架与基于Taalas技术的芯片配对,这意味着一种解耦架构,其中计算密集型的提示处理在GPU上完成,而令牌生成则卸载到基于Taalas的加速器上。

REG AD

REG AD

It’s also possible that AMD could adopt a sort of tick-tock cadence in which customers initially deploy and validate models on Instinct accelerators and, once they’re satisfied with them, transition to Taalas accelerators. We can only speculate at this point, but here’s what AMD’s SVP of AI, Vamsi Boppana, had to say about it in a canned statement:

AMD也有可能采用一种类似嘀嗒的节奏,客户最初在Instinct加速器上部署和验证模型,一旦满意,再过渡到Taalas加速器。目前我们只能推测,但AMD人工智能高级副总裁Vamsi Boppana在一份预先准备好的声明中这样表示:

“AMD is building a full-stack AI platform that gives customers the flexibility to deploy the right compute solutions for every AI workload."

“AMD正在构建一个全栈式AI平台,为客户提供灵活性,以便为每个AI工作负载部署合适的计算解决方案。”

You better really love that model

你最好真的喜欢那个模型

While the tech is blazing fast, if you hadn’t already figured it out, it comes with a pretty substantial downside. Once the chips are deployed you’re stuck with that model. Any change bigger than something like a LoRA adapter is going to require a re-spin of the chips, which is not only expensive but time-consuming.

虽然这项技术速度极快,但如果你还没意识到,它有一个相当大的缺点。一旦芯片部署完毕,你就被那个模型困住了。任何比LoRA适配器更大的更改都需要重新流片,这不仅昂贵而且耗时。

Nearly four years into the AI boom, new models are rolling out on a nearly monthly basis. In order to benefit from Taalas’ tech, AMD’s customers are going to have to be really sure about their choice of models, which will be easier for some than others.

在AI热潮近四年后,新模型几乎每月都在推出。为了从Taalas的技术中受益,AMD的客户必须非常确定他们对模型的选择,这对某些人来说比其他人更容易。

However, if the startup is to be believed, the situation isn’t quite as bad as it sounds. While new models will require a re-spin, it doesn’t require starting over from scratch. Instead, just two layers of metal need to be changed, which is a lot cheaper and less time-consuming.

然而,如果这家初创公司可信的话,情况并没有听起来那么糟糕。虽然新模型需要重新流片,但不需要从头开始。相反,只需更改两层金属,这更便宜且耗时更少。

MORE CONTEXT

更多背景信息

Elon pledges to give Nvidia a virtual monopoly over the stars

埃隆承诺让英伟达在星空中拥有虚拟垄断地位

AMD's AI eggs are in too few baskets, Wall Street worries

华尔街担忧AMD的AI押注过于集中

China turns up the heat with open model blitz as US model makers panic

中国以开放模型闪电战升温,美国模型制造商恐慌

A deep dive into Nvidia's Vera CPU and the Olympus cores that power it

深入探讨英伟达的Vera CPU及其驱动的Olympus核心

With that said, we strongly suspect this tech will largely be deployed by AI model devs, their infrastructure providers, and a handful of inference providers. In an interview with our sibling site The Next Platform in February, the company suggested that etching a model's weights into silicon is 100x less expensive than training a frontier model.

话虽如此,我们强烈怀疑这项技术将主要由AI模型开发者、其基础设施提供商以及少数推理提供商部署。在二月份与我们的姊妹网站The Next Platform的一次采访中,该公司表示,将模型权重蚀刻到硅片上的成本比训练前沿模型便宜100倍。

AMD is certainly in a position to negotiate those deals. OpenAI, Anthropic, and Meta are all major Instinct customers. Given the close working relationship between the model houses and the chip designer, it wouldn't be surprising to see a GPT or Claude deployed on a combination of Taalas and instinct accelerators.

AMD当然有能力谈判这些交易。OpenAI、Anthropic和Meta都是Instinct的主要客户。鉴于模型公司与芯片设计商之间的紧密合作关系,看到GPT或Claude部署在Taalas和Instinct加速器的组合上并不令人惊讶。

REG AD

REG AD

The tech also has implications for model development. One of the ways developers have cut down on hallucinations is by trading time for accuracy. The technique, called test-time scaling, is quite simple in practice, and involves allowing a model to “think” for longer before responding.

这项技术对模型开发也有影响。开发者减少幻觉的方法之一是用时间换取准确性。这种称为测试时扩展的技术在实践中相当简单,涉及允许模型在响应前“思考”更长时间。

One drawback of test-time scaling is that it consumes substantially more tokens, which makes it expensive, and means users have to wait longer for the chatbot, code assistant, or agent to respond. If AMD’s Taalas buy can drive down the cost per token and boost output speeds by 10x or 20x, model devs may opt to extend the reasoning time even further.

测试时扩展的一个缺点是它消耗更多的令牌,这使得成本高昂,并且意味着用户必须等待更长时间才能得到聊天机器人、代码助手或代理的响应。如果AMD收购Taalas能够降低每令牌成本并将输出速度提高10倍或20倍,模型开发者可能会选择进一步延长推理时间。

In any case, we may not have to wait long to see just how Taalas fits into AMD’s broader vision. Subject to regulatory approval, the deal is expected to close in the fourth quarter. ®

无论如何,我们可能不必等待太久就能看到Taalas如何融入AMD的更广泛愿景。在获得监管批准的前提下,该交易预计将在第四季度完成。®

gpu nvidia amd ai semiconductor systems cloud infrastructure month 2026

gpu nvidia amd ai 半导体 系统 云基础设施 月份 2026

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →