跳到主内容
@wquguru
精选70Two Minute Papers(YouTube)模型发布/更新

DeepSeek V4 Pro 发布:开源权重、生成提速 78%

DeepSeek Just Made Closed AI Look Ridiculous

原文
发到 X

Yes, DeepSeek 4 Pro is here. This time, the real one. I know it gets confusing. The previous version was called preview, and this is called 0813. The numbers say it performs better than the much smaller flash version. When building a Rubik's Cube, flash did not completely understand the 3D structure of the object. Lots of missing parts, lots of blackness. But, with the pro, look, much better understanding of structure.

是的,DeepSeek 4 Pro 来了。这次,是真的。我知道这让人困惑。之前的版本叫预览版,而这个叫 0813。数字表明它的性能优于小得多的 flash 版本。在搭建魔方时,flash 没有完全理解物体的 3D 结构。很多缺失部分,很多黑色区域。但是,用 pro,看,对结构的理解好多了。

Then, I was also surprised by this. Holy mother of papers. Look at that. It is inching closer and closer to fable quality. Yep, another challenger appeared, and it gets better. They give all this for us for free, which is absolutely incredible. Okay, but what does it mean for us? The weights are available for free for all of us, but, you know, few have the hardware to host it at home. I'd love to, but I don't have that kind of hardware.

然后,我也对此感到惊讶。天哪,论文之母。看看那个。它越来越接近寓言级质量了。是的,另一个挑战者出现了,而且它变得更好。他们免费给我们所有这些,这绝对令人难以置信。好吧,但这对我们意味着什么?权重对我们所有人免费提供,但是,你知道,很少有人有硬件在家里托管它。我很想,但我没有那种硬件。

Other options include Lambda or using it hosted by DeepSeek themselves. But, they just raised their prices dramatically, about 2 and 1/2 to 5x the previous prices. Now, I bet you can already imagine the clickbait headlines saying, "It's over." I think what they should also say is that DeepSeek has MIT licensed open weights. What does that mean? Well, anyone can run the exact same model at their own price. And, look, they do.

其他选择包括 Lambda 或使用 DeepSeek 自己托管的服务。但是,他们刚刚大幅提高了价格,大约是之前价格的 2.5 到 5 倍。现在,我打赌你已经能想象出那些标题党新闻说“结束了”。我认为他们也应该说的是,DeepSeek 拥有 MIT 许可的开放权重。这意味着什么?嗯,任何人都可以以他们自己的价格运行完全相同的模型。而且,看,他们确实这么做了。

A bunch of hosts available, and they all compete on price. That is amazing for us, and it is very likely to push the Frontier Labs to give us fellow scholars something even better, and quickly. And all this improvement comes from the same architecture. But, how is that even possible? The model structure is the same, yet it is massively better than the preview was less than 4 months ago. So, how? Dear fellow scholars, this is Two Minute Papers with Dr.

有一堆可用的托管服务,而且它们都在价格上竞争。这对我们来说太棒了,而且很可能会推动前沿实验室给我们这些学者一些更好的东西,而且要快。所有这些改进都来自相同的架构。但是,这怎么可能呢?模型结构相同,却比不到 4 个月前的预览版好得多。那么,怎么做到的?亲爱的学者们,这里是《两分钟论文》节目,我是 Dr.

Károly Zsolnai-Fehér. Once again, a lot of the magic happens after pre-training. During post-training, DeepSeek creates several specialist models for mathematics, coding, and agentic work. Now, we have to stop here for a moment. People confuse these with the experts in mixture of experts. That's not quite the same. Those are little pieces within one neural network. These are not. These are separately trained model checkpoints.

Károly Zsolnai-Fehér。再一次,很多魔法发生在预训练之后。在后训练期间,DeepSeek 创建了几个专家模型,用于数学、编码和智能体工作。现在,我们必须在这里停一下。人们把这些和混合专家中的专家混淆了。那不完全一样。那些是单个神经网络内部的小部分。这些不是。这些是单独训练的模型检查点。

Okay, so what then? Then comes distillation. Yeah. They take more than 10 of these specialist teachers and train one final model to absorb their abilities. So, the student model says, "This is what I would do." Then the teacher says, "Well, this is what I would have done." Then the student adapts its brain [clears throat] to be more like its teacher. Do it with 10 teachers and you see that the student indeed improves like crazy.

好的,那接下来呢?接下来是蒸馏。是的。他们用超过10个这样的专家教师模型,训练一个最终的学生模型来吸收它们的能力。学生模型说:“我会这样做。”然后教师模型说:“嗯,我会那样做。”于是学生模型调整自己的“大脑”(清嗓子)变得更像它的老师。用10个老师这样做,你会发现学生模型确实进步神速。

They also added this part to it. Instead of just predicting one token at a time, it drafts several tokens ahead. It does it much better than previous techniques. And hold on to your papers, fellow scholars, because DeepSeek reports up to 78% faster generation for V4 Pro. Real, measurable speed up in real use that you get right now and benefit from it. And here is something absolutely insane. This was a research paper, let's see, 6 weeks ago.

他们还加入了这一部分。不只是每次预测一个词元,而是提前草拟多个词元。这比之前的技术做得更好。各位学者,请抓紧你们的论文,因为DeepSeek报告称V4 Pro的生成速度提升了高达78%。这是真实、可衡量的速度提升,你现在就能用上并从中受益。接下来是绝对疯狂的事情。这是一篇研究论文,让我们看看,6周前发布的。

And now everyone is using it. Let me say it again, a research paper only 6 weeks ago. One of the best papers of the year. And it is coming alive right in our hands, for free. Incredible. Full breakdown video in the description. And don't forget, we own and can run the weights. No one downgrades us to a different model if we type the wrong keyword. No games. That is incredible. Even if I can't run it at home, which I would love to do.

而现在每个人都在使用它。让我再说一遍,仅仅6周前的研究论文。年度最佳论文之一。它正在我们手中活起来,而且是免费的。太不可思议了。完整解析视频在描述里。别忘了,我们拥有权重并且可以运行它们。如果我们输入了错误的关键词,没有人会给我们降级到另一个模型。没有套路。这太棒了。即使我无法在家运行它,虽然我很想。

But, there are options. What a time to be alive. This is open science and open research at its best. And it's important that we talk about it. Why? Because the future belongs to those who understand it. Use DeepSeek and use DeepSpark. Take advantage of them. Oh, and I plan to talk about DeepSeek's incredible no agent harness as well. Novel design, really powerful. If you're interested, consider subscribing and hitting the bell.

但是,有选择。活在这个时代真好。这是开放科学和开放研究的极致体现。我们谈论它很重要。为什么?因为未来属于那些理解它的人。使用DeepSeek和使用DeepSpark。好好利用它们。哦,我还计划谈谈DeepSeek那不可思议的无代理(agent)框架。新颖的设计,非常强大。如果你感兴趣,考虑订阅并点击铃铛。

I use Lambda to reproduce AI research papers often in minutes. It's also great to train your own models or fine-tune an existing one. Run inference or text-to-image or video, easy-peasy. Running a DeepSeek chatbot or agent, super fast, super reliable. Lambda gives you powerful Nvidia GPUs to run your own experiments. I test ideas from the papers I cover and moments later, results. Love it. Seriously, try it out now at lambda.ai/papers.

我使用Lambda来复现AI研究论文,通常几分钟就能完成。它也非常适合训练你自己的模型或微调现有模型。运行推理或文本到图像、文本到视频,轻松搞定。运行DeepSeek聊天机器人或代理,超快、超可靠。Lambda为你提供强大的Nvidia GPU来运行你自己的实验。我测试我报道的论文中的想法,片刻之后,结果就出来了。太喜欢了。说真的,现在就去lambda.ai/papers试试吧。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近