自建数据中心:算力成本核算与训练推理分层策略
How to Build Your Own Data Center & Why Every Startup Should Do It
给出了具体的算力持有vs租赁成本对比数字,以及训练/推理分层的实操框架,适合有AI业务且规模达到一定阈值的团队参考其单位经济模型。
You said 11 raps kind of leaprogged you.
你说有11次机会让你突然开窍了。
Is that on you?
这算是你的责任吗?
Yeah, 100% it's on me. It's the biggest strategic mistake I made in the history of Speedify.
是的,百分之百是我的责任。这是我在Speedify历史上犯下的最大的战略错误。
How do you reflect on that?
你是如何反思这一点的?
So today is a real freaking discussion. Cliff Whitesman, founder and CEO at Speechify, one of the fastest growing text to speech startups in the world on the show.
所以今天是一场非常棒的讨论。Cliff Whitesman是Speechify的创始人兼CEO,Speechify是全球增长最快的文本转语音初创公司之一,他来到了我们的节目。
The best way to lose is not to be in the race. Be in the race. You don't want to be a fat manager who is like a general sitting in the back saying, "Take that hill." You want to be the warrior who runs up with their sword and engages the enemy first.
输掉比赛最好的方式就是根本不参加比赛。要参加比赛。你不想成为一个像将军一样坐在后面说‘拿下那座山丘’的肥胖经理。你想成为那个挥舞着剑冲上去、率先与敌人交战的战士。
Ready to go.
准备就绪。
Cliff, it is so good to have you back in the studio, dude. I I was looking forward to this one cuz when I was writing it up, it's a very different thread of conversation to how I'd normally go. And so, thank you so much for joining me again today, dude.
Cliff,很高兴你再次回到演播室,伙计。我一直很期待这次对话,因为当我在撰写提纲时,我发现这是一条与我平时风格截然不同的谈话线索。所以,非常感谢你今天再次加入我,伙计。
My pleasure. Glad to be here as always. Now, I wanted to start with you're spending tens of millions of dollars on Nvidia GPUs and you're paying an additional $100,000 per GPU to receive them 4 months early. Why? Like, what do you know that the market doesn't know? So, in 2022, we bought a huge rack of GPUs from Nvidia. And the reason we bought them is for training, right? We have a bunch of models. The newest Speedify Simba 3.2 2 model is ranked number one in the world for quality um above all the frontier labs 10x more affordable and stuff like 11 labs and we used to rent GPUs and we found that engineers at speechify would be parsimmonious with how they use the GPUs cuz they were like oh my god I'm costing the company tens of thousands of dollars like I don't want to do that and the analogy my brother and I came up with is imagine you're Michael Jordan and you want to be in the NBA it's the only thing you care about and you need to pay $20 an hour just to train in a basketball center well that sucks you want one that you can go to whenever ever you want to.
这是我的荣幸。很高兴能像往常一样来到这里。现在,我想从你开始在Nvidia GPU上花费数千万美元说起,而且你每块GPU还要额外支付10万美元以提前4个月收到货。为什么?你知道什么市场还不知道的吗?嗯,在2022年,我们从Nvidia购买了一整排GPU。我们购买它们的原因是为了训练,对吧?我们有一堆模型。最新的Speedify Simba 3.2 2模型在质量上排名世界第一,优于所有前沿实验室,价格实惠10倍等等,比如11 labs。我们过去租用GPU,发现Speechify的工程师在使用GPU时会比较吝啬,因为他们觉得‘天哪,我正在让公司花费成千上万美金’,所以他们不想那样做。我和我兄弟想到的类比是:想象一下你是迈克尔·乔丹,你想进入NBA,这是你唯一关心的事情,但你却需要每小时支付20美元才能在篮球中心训练。那太糟糕了,你想要一个可以随时去的地方。
In fact, you want a hoop in your house. And so our initial idea was we want a hoop in our house. And so we bought a bunch of our own GPUs. And that deal ended up being really good for us. And we ended up training really good models. So with time we invested more and more and more and more. So that's the first part. The second part is actually how the economics work out. So if you look at it, the transformer was invented inside of Google in 2017.
事实上,你希望家里有个篮筐。所以我们最初的想法是我们想要在家里装个篮筐。于是我们买了一些自己的GPU。这笔交易对我们来说非常划算。最终我们训练出了非常好的模型。因此随着时间的推移,我们投入越来越多。这是第一部分。第二部分是经济账是怎么算的。如果你看看它,Transformer是在2017年在Google内部发明的。
Nvidia came out with A100 GPUs in 2019. Shortly after they came out with H100 GPUs, right? The original Chetchup PT was trained on A100s. And then they came out with Blackwells. So then B200s, B300's, and now they came out with Reubins, which is the GPUs that Elon is sending to space. And they're like liquid cool cooled. They're very, very cool. And we're like, okay, huh. One, every class of GPU is more affordable per one trillion flops, right?
英伟达在2019年推出了A100 GPU。不久之后他们又推出了H100 GPU,对吧?最初的Chetchup PT是在A100上训练的。然后他们推出了Blackwell架构。所以接着是B200、B300,现在他们又推出了Rubin,也就是埃隆·马斯克发送到太空的那种GPU。它们采用液冷散热,非常非常酷。我们心想,好吧,嗯。第一,每一代GPU每万亿次浮点运算的成本都更低,对吧?
So, a flop is addition, subtraction, multiplication, any mathematical operation, and you measure them in how many trillion of operations happen per second in a GPU. And so, they're more affordable as it relates to this. Um, if I was to buy an H100 for, let's say, $30,000, that's how much the kind of a single card would cost. If I wanted to rent an H100 for one hour spot instance from GCP, it could cost me $5. If I rented it from like, you know, Azure or AWS, maybe it'll cost me $3.5 per hour.
那么,一次flop(浮点运算)指的是加法、减法、乘法或任何数学运算,我们通过GPU每秒能执行多少万亿次操作来衡量它。因此,就这一点而言,它们的性价比更高。嗯,如果我花3万美元买一张H100,这就是单张卡的大致价格。如果我想从GCP租用一个H100的按小时计费的实例,可能需要5美元。如果我从Azure或AWS租用,可能每小时只需3.5美元。
So, if I multiply that times 24 hours and then times 365 days in a year, I'm actually going to end up paying $35,000 to $50,000 to rent that GPU for one year, but I could buy it for $30,000. So, it's 1.5x the cost of owning the hardware to rent the hardware for a year. Now, the hardware is typically um warrantied for 3 years to work properly, but it'll keep working up to the warranty for I imagine, I don't know, 10 years.
所以,如果我把这个价格乘以24小时,再乘以一年的365天,我实际上需要支付3.5万到5万美元来租用该GPU一年,但我可以花3万美元买下它。因此,租用硬件一年的成本是拥有硬件成本的1.5倍。现在,硬件通常保修3年以确保正常工作,但我想即使过了保修期,它也能继续工作长达10年。
So the math just maths where it makes way more sense to buy them. The other big part is if you want to do large scale training like we do, you need the memory to be colllocated with a large cluster of GPUs. I can't just rent from Google or Microsoft or even B 10 and run the size of training that I want because I need a gigantic memory card next to it with all of my data that all the GPUs are accessing. So that's why we first started buying them.
所以算下来,购买显然更划算。另一个重要部分是,如果你想进行像我们这样的大规模训练,你需要内存与大型GPU集群位于同一位置。我不能仅仅从Google、Microsoft甚至B10租用,然后运行我想要规模的训练,因为我需要在旁边配备一块巨大的内存卡,里面存储着所有GPU都在访问的数据。这就是为什么我们最初开始购买它们的原因。
The next thing that we found is actually if you run open- source models for coding, you could pay anthropic and then you know you're paying for all the tokens and the fact that you're doing the the the branded right fable one or you can run an open source model and instead of running it on a spot instance from Azure or anyone else you run it on your own hardware and then you're paying a fraction of a fraction of a cent per token and so for all those reasons it made a ton of sense but we can go into all the depth that you want.
我们要发现的下一件事是,如果你运行用于代码生成的开源模型,你可以向Anthropic付费,你知道你是在为所有的token付费,而且你是在使用那个品牌化的Fable One;或者你可以运行一个开源模型,而不是在Azure或其他人的按小时计费实例上运行它,而是在你自己的硬件上运行它,然后你每个token支付的费用只有几分之一美分。出于所有这些原因,这非常有意义,但我们可以根据你的需求深入探讨每一个细节。
I just want to dig in. The first thought that I have is I completely understand the rationale there but chips depreciate. You have chip cycles and they are accelerating. We are seeing newer and newer chips being created. We're seeing specialization within chips by buying your locking yourself in so to speak to one chip architecture. How do you think about that? At speechify, we still use K80s for a lot of specific operations for inference and we use older models of GPUs constantly.
我只是想深入探讨一下。我首先想到的是,我完全理解其中的逻辑,但芯片会贬值。存在芯片周期,而且这一周期正在加速。我们看到越来越多的新芯片被创造出来。我们还看到芯片内部出现专业化趋势,比如通过锁定(locking yourself in)将自己绑定在一种芯片架构上。你是如何看待这一点的?在 Speechify,我们仍然在许多特定的推理操作中继续使用 K80,并且不断使用旧型号的 GPU。
Um, and there's essentially a difference between when you do inference and when you do training. For training, I'm like, okay, I have this hypothesis. I want to know the answer to this hypothesis as soon as possible. Like every minute that it doesn't come out, I'm in competition with everybody else. And so having a GPU architecture that is much faster by orders of magnitude is a huge advantage. But if you go speech to text or text to speech with speechify, I can afford to give you a lower quality GPU and it'll give you what you need still in, you know, 100 milliseconds.
嗯,推理和训练之间本质上是有区别的。对于训练来说,我会想:好吧,我有一个假设,我希望尽快知道这个假设的答案。就像每多一分钟得不到结果,我就在与其他人竞争。因此,拥有速度快几个数量级的 GPU 架构是一个巨大的优势。但是,如果你进行 Speechify 的语音转文字或文字转语音处理,我可以给你提供较低质量的 GPU,它仍然能在大约 100 毫秒内为你提供所需的结果。
So it's like totally good. And so I can always use these older GPU models for inference. That's number one. Number two, we have so many experiments that were running at every single point in time. Not all of them need to run on like the newest hardware. So, the analogy I always give, let's say you bought an iPhone back in 2011 and it's an iPhone 3G and then you bought another iPhone and another iPhone and another iPhone.
所以这完全没问题。因此,我总是可以使用这些旧型号的 GPU 进行推理。这是第一点。第二点,我们在每个时间点都有大量的实验在运行。并非所有实验都需要在最新硬件上运行。所以我常举这样一个例子:假设你在 2011 年买了一部 iPhone 3G,然后又买了一部又一部又一部 iPhone。
You could have a drawer in your house with like five iPhones that are collecting dust cuz you can only use one iPhone at a time, but if I own a 100,000 GPUs, I'm still going to use all of them at the same time. And so, I'm not losing anything by having more GPUs because not only do I own a bunch, I still rent from the hyperscalers all the time. And I rent both dedicated instances that I prepaid for and I rent spot instances.
你家里可能有一个抽屉,里面放着五部积灰的 iPhone,因为你一次只能使用一部 iPhone;但如果我拥有 10 万台 GPU,我仍然会同时使用它们。因此,拥有更多 GPU 并不会让我损失什么,因为我不仅拥有一大批 GPU,还一直向超大规模云服务商租用资源。我既租用预先付费的专用实例,也租用抢占式实例。
For example, more people use speify in September because everybody goes back to school. So I need to like level out the load. And so the parts of that load that I know for sure I'm always going to use whether it be training or it be inference, I might as well just own it. And then on top of that is also the case that I have so many other friends who are running training and running inference. I can always rent it out to other people if I have excess capacity, which I don't expect to have.
例如,9 月份更多人使用 Speechify,因为大家都返校了。所以我需要平衡负载。对于那些我确定无论何时都会使用的负载部分,无论是训练还是推理,我干脆直接拥有它们。此外,还有这样的情况:我有许多其他朋友也在运行训练和推理任务。如果我有剩余容量(尽管我不期望有),我总可以将它们出租给别人。
But like every once in a while you have an interesting situation. So for all those reasons, it just makes mathematical financial sense. Lastly, if you have excess capital really, you either stick it on the bank or you buy a bond, right? Like the best long year, the best bond you can buy long tail, I don't know, will yield you like 5%. Or you can buy a GPU and because renting it would cost me 1.5x buying it for the year, the return is like way higher.
但偶尔你会遇到一些有趣的情况。因此,出于所有这些原因,这在数学和财务上都是合理的。最后,如果你真的有闲置资金,你要么把它存进银行,要么买债券,对吧?就像过去表现最好的一年,你能买到的长期债券,我不知道,收益率大概只有5%。或者你可以买一块GPU,因为租用它的成本是我购买它一年的1.5倍,所以回报率要高得多。
So how many GPUs do you buy then?
那么你要买多少块GPU呢?
So let's talk about Reuben's for example. So Reubins come in the form of 72 cards in one rack. So we'll buy multiple racks of Reubins and then on top of that we'll buy B300s which are like the newest form of Blackwells uh because we can get them earlier. And then the same thing like you know when we bought our f
让我们以Rubin为例来谈谈。Rubin的形式是一个机架中有72张卡。所以我们会购买多个机架的Rubin,此外我们还会购买B300,这是Blackwell系列最新的形式,呃,因为我们能更早地拿到它们。然后情况也是一样的,比如当我们购买我们的f
原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力