跳到主内容
@wquguru
精选80Two Minute Papers(YouTube)模型发布/更新

英伟达开源Nemotron 3 Ultra:550B参数,速度极快

NVIDIA's New Free AI - A Gift To Humanity

原文
发到 X

This AI is not NeMoTron 3 Super. No. This is NeMoTron 3 Ultra, Nvidia's newest free and open AI model. And I've been delighted, disappointed, and confused by it. But I think I got it now. You see, you can look at the benchmarks all you want, but we are fellow scholars here. We don't just believe stuff. We test it for ourselves. That is the way of the scholar. So, I had an early look at it and ran some of my experiments day and night.

First impression is that it is incredibly fast, blazing fast. Love that. But then, my coding experiments did not go that well. When I asked it to write a light simulation program, this is my original area of research, and I get a black screen, nothing. When I asked it to fix it, it does a bunch of things, and same. Then I said, "Okay, let's debug this by hand." It had some mistakes. After fixing that, well, we get something.

But maybe it's a scene that does not work at all. Other even smaller systems can do this task with relative ease. And the other thing is, goodness, it wrote up more than a thousand lines of code. Oof. You don't need that much. My handwritten solution from my research is about 250 lines and renders this scene. Fully open source, free for everyone, forever. Now, let's write a real-time strategy game. Yes. No. Black screen again.

Almost. We got a square. But if you ask [clears throat] DeepSeek for flash with the same prompt, you get something really cool, but not here. So, what is going on here? Well, I went back and forth with Nvidia and reported some of the issues, and later there were some improvements. But still, this kind of coding is not something I would personally use this for. So, I said, you know, maybe let's not use this AI. But then I thought, wait, it is super fast and probably good at other things.

So, I gave it agentic things. Fixing broken installations on my machine from the terminal, excellent. Whipping up quick experiments, organizing files, excellent. Super fast. And over time, I found myself reaching out to it more and more. And I [clears throat] found it to be useful basically for everything other than challenging coding tasks. Now that is excellent because this might be the openest AI model ever. Weights are open, the research paper on how it was made is open, training data and recipes are being released at least for the redistributable parts.

Now that is pretty crazy. Now, hold on to your papers, fellow scholars, because it gets even better. Licensing, super important question, very overlooked. We are always hoping for Apache 2.0. This is the do what ever you want license. For me, this is 10 out of 10. Now, Nvidia started publishing their models under their own proprietary license, which I would rate seven out of 10. Derivative works and commercial use is fine.

On the other hand, it needs a bit of attribution and a little stricter on patent grants. Now, this has the Open MDW license. This is basically Apache 2.0 tailored for machine learning weights. This is absolutely fantastic news, glorious. I think this might be a nine out of 10. Maybe as close to 10 out of 10 as you can get from a big company like Nvidia. Allows basically everything, but less battle tested and my understanding is that if you sue claiming this model infringes your rights, you lose the license.

Huge improvement, double thumbs up. Thank you. Now, can you run it yourself? Hmm, um yes and no. Yes, because completely open. Download it, it is yours forever. No limits, no funny business. However, no, because I would love to run it locally, too, but it's huge. 550 billion parameters. You need hundreds of gigabytes of GPU memory for that. This is why I will probably use it on Lambda. Also, 1 million token long context window.

Ah. Great. Have a larger code base with a bug hiding somewhere? No worries. Massive books? Easy. Okay, how about images and videos? Well, it does not have vision capabilities. Not multimodal. Text only. Oh. Man. How much I would love a multimodal version of this. Goodness. Please. Okay, and I also had a realization. You don't need one model to do everything. You need a roster of models that cover your use cases. For instance, I can't add vision capabilities to Nimatron 3 Ultra, but I can both Gemma 4 to it with a screwdriver.

It's like a seeing eye dog guiding a smarter blind man along. It is hilarious, and it kind of works. Kind of. So, we finally have more competition in the open AI model space, and that is glorious. So, how does it work? Well, one trick is that it is huge, but not all of it runs at once. 550 billion parameters total, but only about 10% of that is active per token. These are specialist mini brains that are being activated at a time.

We call that mixture of experts. But, you wise fellow scholars know that already. So, what else? Now, they also use Mamba layers. Why Mamba? Is this like a snake or like the fruity chew? I don't know. I don't even know why I brought this up. [laughter] So, what do these do? Well, traditional AI systems have a bit of a memory problem. They work like a student who constantly rereads the textbook over and over again when they are given a question.

But, memory is precious. So, instead, read the book only once and take highly compressed notes. So, this kind of memory remembers important details about the conversation. However, it is also smart enough to throw away the filler words. Thus, this system can process massive amounts of data efficiently. It also uses low precision numbers, so you have to do less number crunching when running this. They call it NVFP4, and this doesn't rely on predicting tokens one by one.

No, it has multiple heads that draft multiple future tokens at the same time. Once again, many things that make it blazing fast. And, we get all of this for free forever. What a time to be alive. Thank you to everyone who worked on this, and absolutely everyone everywhere who is working on open-source projects and open models. You are all heroes. And, look, this system is great, but it could be tiny. It could be bad, ugly.

I don't care. As long as it is open science and open models, it pushes humanity forward. Thank you. What a time to be alive. Here you see me running the full DeepSeek AI model through Lambda GPU Cloud. 671 billion parameters running super fast and super reliably. This is insane. I love it, and I use it on a regular basis. Lambda provides you with powerful Nvidia GPUs to run your own chatbots and experiments. Seriously, try it out now at lambda.ai/papers or click the link in the description.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近