英伟达挑战者d-Matrix:AI推理芯片比GPU快10倍
This Nvidia Challenger Says Its AI Chip Is 10x Faster Than A GPU
There's a new Nvidia Challenger, d-Matrix, making a very different kind of chip it says is far faster at AI inference than Nvidia's standalone GPUs. And we are reporting here first that it's now in full production. Our solution solves the problem that GPUs cannot. GPUs cannot, Trainium chips cannot. So they are in high demand. The new chip, called Corsair, is made at TSMC in Taiwan on six nanometer node and will begin shipping later this month.
If you take out the sheet metal on this product, you'll see, you know, collection of four chips inside. This is something you can plug into a server. d-Matrix won't disclose specific customers yet, but says it already has commitments from high profile hyperscalers, neoclouds and leading AI labs eager to get their hands on as much compute as possible. d-Matrix pairs the custom accelerator with graphics processing units in a server rack that it says produces tokens ten times faster than GPUs can on their own.
It also says Corsair is three times cheaper, up to five times more energy efficient, ideal for inference workloads where speed is the priority like chat bots, video generation and, of course, Agentic AI. The California based chip maker was founded in 2019 by Sid Sheth. He says Corsair bypasses the memory bottleneck that's panicking the AI industry. You're not running into a choke point around D ram with our product, because our product doesn't really rely on D ram to be successful.
Unlike GPUs, which rely on large amounts of high bandwidth memory, that's in short supply from makers SK Hynix, Samsung, and Micron. d-Matrix's custom accelerator uses a different kind of memory. S ram integrated directly into the chip to greatly increase how quickly that memory can be accessed. This alternative approach to memory is a model that's brought two other once small chip makers major gains this year. Cerebras, which had a blockbuster IPO at nearly $100 billion market cap in May over excitement around its memory first dinner plate size chips, and Groq.
In December, Nvidia bought Groqs assets for $20 billion in its biggest purchase to date. Then at GTC in March, Nvidia unveiled an entire new line of chips made with Groqs technology called Language Processing Units, or LPUs. Additional new tiers of inference performance for token generation the world's never seen. So this is it. This is Groq. Groq and Cerebras were able to kind of take the S ram and put it onto the same chip.
We did the same thing and now it's the same chip. And so you're spending less time and energy moving the data back and forth. Nvidia is also reportedly preparing a version of its Groq chips to sell in China, where U.S. export controls prohibit the sale of its most advanced flagship GPUs. So could the China market be a big opportunity for the new d-Matrix chips too? Over time, we might be able to ship product to China to just run AI models more efficiently, because we don't give them the ability to train those models on our chip.
In November, d-Matrix announced a $275 million funding round, putting it at a $2 billion valuation. Microsoft was one of the investors through its M12 venture arm. That's notable because of Microsoft's own chip ambitions. From its own Maya chips for AI inference to new PC processors built with Nvidia and an in-house chip for quantum announced last week. d-Matrix also teamed up with Arista, Broadcom and Supermicro to build a full rack scale system for deploying its chips in AI data centers.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力