NVIDIA新AI学会跑酷:仅需30秒人类视频,模仿与目标兼顾
NVIDIA's AI Learns Why Copying Humans Isn't Enough
This is one of the most fun papers I've read in a while. You will love this one. So, what if you knew every athletic move perfectly, but had no idea when to use which? A perfect vault is useless if the table moves away. This is the strange problem of teaching parkour to a virtual human. Systems that copy us are beautiful, but brittle, repetitive. Systems that chase goals adapt, but often stop moving like us. But this new way I work says, "Enough is enough.
Let's do both." It promises to remember exactly how humans move, yet is able to break the script when needed. Hmm. And get this, here's the crazy part, the training data. They say that all you have is only 30 seconds of internet parkour to pull this off. That, fellow scholars, sounds impossible. You see, previous techniques were not up to the task. This guy moves like a drunkard. Not to worry, try again. And
[laughter]
Yep, get some rest. Next. Now this guy looks a bit more promising. Wait a second, that's cheating. You are supposed to get through the level, not avoid it, mister. That is cheating. Get out of here. Previous technique number three. Ooh, now we're talking. Look at that. Oof, what an insane research field. This is so much fun. Okay, so can this new technique solve this because the previous ones couldn't? Let's see if it can solve the level.
I don't think so. Dear fellow scholars, this is Two Minute Papers with Dr. Kairan and I here. Now hold on to your papers, fellow scholars, and let's see. That's a strong start, very athletic. Hmm. [clears throat] Ooh, look at that. Solving the level like a ninja. It really adapts and composes new skills to get through the level. Okay, so we are experienced fellow scholars here. So, we ask, was this the only level it can finish or is this the real deal?
Can it solve longer levels? Let's see. Oh my, we are off to a good start. Still going. Yes, incredible. Good job. Ouch. Okay. Get well soon, little AI. Now, after taking a trip to the AI hospital, let's try a variety of levels at the same time. Ooh, this is amazing. Love it. Now, let's make it harder. A level with unseen obstacle arrangements. And oh my, still doing amazingly well. And the final jump. Now, that is a sight to behold.
And it gets better because I will show you my favorite move and it's kind of the best. Okay. So, how on earth did they do that? Well, they put one character into two classrooms. In the first, it learns to copy human movement. This produces beautiful human-like motions, but it is limited by the training data. It cannot go beyond what it has seen. And remember, that is only 19 clips containing 30 seconds of parkour. Very little data.
In the second classroom, it learns to solve new obstacle courses. And from this point on, the same AI controller learns in both classrooms simultaneously. Do not forget about the part where both classrooms give it body, obstacle, and destination information. Super important. And here comes the most beautiful part. Hmm, it's quite the formula, but we'll break it down. I'll try my best. My best way to imagine this is that now we train a judge to minimize three kinds of error.
Here are real human movements together with the obstacles around them. The judge should call these real. Then here the movements are produced by the AI. The judge should call these fake. Now we use this judge to train our AI athlete. If it moves like a clumsy AI, it gets a poor score. Get out of here. But as its movements become more human-like and appropriate for the obstacle, it gets a treat, a good score. And the judge and athlete keep improving together.
The athlete gets better at fooling the judge while the judge gets better at spotting artificial movements. This is absolutely beautiful. Finally, this last term keeps the judge stable. So tiny changes in a movement do not completely change its decision. And this is what leads to these absolutely stunning results. But it not perfect. It offers lower tracking error than its predecessors, but at a cost of a bit of a hit to the success rate.
This trade-off is among the weaker spots, I think. So I would chalk this up as a limitation. Longer levels still only have about 40% success. The other limitation is that unnatural recovery motions are also possible. Now, as promised, here is my favorite move in the sequence. Woo! Wow! This is incredible. What a time to be alive. But here is the most important thing. The research paper is freely available. Free knowledge for all of us and potentially code later, too.
So all of this knowledge is out there for all of us for free and nobody is talking about it. This is driving me mad. This is why I keep making these videos. This is what I live for. Every research paper is a gift to humanity and I want to make sure that humanity receives this gift. I use Lambda to reproduce AI research papers often in minutes. It's also great to train your own models or fine-tune an existing one. Run inference or text to image or video, easy peasy.
Running a deep seek chatbot or agent, super fast, super reliable. Lambda gives you powerful Nvidia GPUs to run your own experiments. I test ideas from the papers I cover and moments later, results. Love it. Seriously, try it out now at lambda.ai/papers.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力