AI教母李飞飞谈世界模型:超越ChatGPT的下一个前沿
What's Beyond ChatGPT? The 'Godmother of AI' Has a Plan | The Circuit
I hear you love walk and talks. Yes, because we sit too much. Yeah. And also our company is just next to this park, so it's great. Is it all AI talk all the time here? I don't know other people's talk, but I talk AI all the time. But that's the only thing I know. Fei-Fei Li is a giant in the field of computer vision, a branch of AI dedicated to teaching machines to see. She made a breakthrough with ImageNet, a visual library that sparked the modern AI revolution, earning her the title, 'Godmother of AI'.
听说你喜欢边走边聊。是的,因为我们坐得太久了。是啊。而且我们公司就在这个公园旁边,所以很方便。这里是不是一直都是在聊AI?我不知道别人聊什么,但我一直在聊AI。但那是我唯一懂的东西。李飞飞是计算机视觉领域的巨擘,计算机视觉是人工智能的一个分支,致力于教机器看懂世界。她通过ImageNet取得了突破,这是一个视觉库,引发了现代人工智能革命,为她赢得了“AI教母”的称号。
And she's been a pioneer at AI's center ever since, as a Stanford professor, a Google executive, an advisor to US presidents, and a relentless advocate for AI that serves people. I feel the leaders of AI are not talking enough about empowering humanity. They're talking too much about replacing humanity, replacing jobs. Now she's adding startup co-founder to her star resume. What's it like running your own shop? I feel like a tiger mom.
此后,她一直是人工智能领域的先驱,担任斯坦福大学教授、谷歌高管、美国总统顾问,并坚定不移地倡导以人为本的人工智能。我觉得AI的领导者们对赋能人类的讨论还不够多。他们谈论更多的是取代人类、取代工作。现在,她又在自己的明星履历上增添了创业公司联合创始人的头衔。经营自己的公司感觉如何?我感觉自己像个虎妈。
While the rest of the AI world doubles down on large language models like ChatGPT and Claude, Li is betting on a new frontier, world models. She's building AI that aims to predict what happens next in the real world, not just the next word in a sentence. So what can world models do ultimately that LLMs will never be able to? Can words put down fires? Can words cook an omelet? Li and her startup, World Labs, is competing in a crowded field of rivals, all making the same bet that world models are the next big leap in AI.
当AI界的其他人都专注于像ChatGPT和Claude这样的大型语言模型时,李飞飞押注于一个新的前沿领域——世界模型。她正在构建旨在预测现实世界中接下来会发生什么的人工智能,而不仅仅是句子中的下一个词。那么,世界模型最终能做什么是LLM永远做不到的?文字能灭火吗?文字能煎蛋卷吗?李飞飞和她的初创公司World Labs正在与众多竞争对手竞争,这些对手都押注世界模型是AI的下一次重大飞跃。
Did you ever imagine AI would be this big? Have I imagined the entire civilization is going to be redefined by AI? The cultural shifts, the geopolitical landscape? I would not have imagined that. Before she became a leading AI scientist, Li was just a curious kid looking at the world. What was young Fei-Fei like as a kid? It was '80s and '90s in China. My city was Chengdu. My dad is quite a curious soul about nature, so I got quite a dosage of chasing after bugs and, and running after puppies and going to the mountains to, to draw and sketch with my dad.
你有没有想过AI会变得如此强大?我有没有想过整个文明会被AI重新定义?文化变迁、地缘政治格局?我没想到过。在成为顶尖AI科学家之前,李飞飞只是一个好奇的孩子,观察着这个世界。小时候的飞飞是什么样的?那是80年代和90年代的中国。我的城市是成都。我爸爸对大自然充满好奇,所以我经常跟着他追虫子、追小狗,还和他一起去山上画画写生。
And, you know, reading. Cutting your hair short, loving aerospace and physics. Did you see yourself as a rebel? As a kid, right? You're preteen. You almost have to be a rebel just because you want to make a statement for yourself. But the love of nature and science and physics, the love of aerospace and outer space, those are part of my identity. You emigrated from China to the US in - New Jersey. In the 90s. Yeah. What was it like coming here as a teenager back then?
而且,你知道,阅读。剪短发,热爱航空航天和物理。你把自己看作叛逆者吗?小时候,对吧?你还没到青春期。你几乎必须成为叛逆者,仅仅因为你想为自己发声。但对自然、科学和物理的热爱,对航空航天和外太空的热爱,这些是我身份的一部分。你在90年代从中国移民到美国——新泽西。那时候作为青少年来到这里是什么感觉?
That was tough. Shifting the whole thing, culture, language, life as a teenager is perhaps one of the hardest times because you're finding who you are and then suddenly you don't even know the world, right? Li learned English, finished high school, studied physics at Princeton while running the family dry cleaning business on weekends, and followed her dreams to Caltech, where she earned her PhD and became a computer scientist.
那很艰难。整个转换,文化、语言,作为青少年的生活也许是最困难的时期之一,因为你在寻找自己是谁,然后突然你甚至不了解这个世界,对吧?李学会了英语,完成了高中学业,在普林斯顿学习物理,同时周末经营家里的干洗店,并追随梦想去了加州理工学院,在那里获得了博士学位,成为了一名计算机科学家。
You've been studying artificial intelligence since then. It was super niche. Yes. What was the field like back then? That was a beautiful period. It was really pure curiosity because there was no money, there was no power, there was no fanfare, there was no hype. We were there just the same as scientists look into the sky, go down in the ocean. I was deeply curious. How does AI see and how is it different from how you or I see?
从那以后你一直在研究人工智能。那是一个非常小众的领域。是的。那时候这个领域是什么样的?那是一段美好的时期。真的是纯粹的好奇心,因为没有钱,没有权力,没有喧嚣,没有炒作。我们就在那里,就像科学家仰望天空,潜入海洋一样。我深感好奇。人工智能是如何看待世界的,它与你我看待世界的方式有何不同?
We use our eyes to drive the sensory system to turn that into signals in the brain, and the brain does computing. It's not that different for computers. We need to learn the pattern of the world and understand objects, colors, environments. A lot of seeing is preparing humans to act. You wake up this morning, you hug your kids, you go downstairs, prepare some breakfast, grab your key, drive to your work. Everything I've described so far, you've done that with your spatial intelligence.
我们用眼睛驱动感觉系统,将其转化为大脑中的信号,大脑进行计算。对计算机来说,这并没有太大不同。我们需要学习世界的模式,理解物体、颜色、环境。很多视觉是为了让人类准备行动。你今天早上醒来,拥抱孩子,下楼,准备早餐,拿钥匙,开车去上班。到目前为止我描述的一切,你都是通过空间智能完成的。
So it's got to be hard to get computers to do that. Yes. Li was working on this back in 2006. She built ImageNet, a massive catalog of 14 million pictures in over 21,000 categories, the largest assembled at that time, because Li saw that algorithms needed data to get smarter. So what was the spark for ImageNet? Nobody was paying attention to data. My students and I had that epiphany that learning needs to be driven by data.
所以让计算机做到这一点一定很难。是的。李在2006年就致力于此。她构建了ImageNet,一个包含超过21,000个类别中1400万张图片的庞大目录,这是当时最大的汇编,因为李看到算法需要数据才能变得更聪明。那么ImageNet的灵感火花是什么?没有人关注数据。我和我的学生们有了那个顿悟:学习需要由数据驱动。
In 2010, Li turned ImageNet into a competition. Who could build the best algorithm to correctly identify the most images? In 2012, Geoffrey Hinton's University of Toronto team entered with AlexNet, an algorithm powered by NVIDIA Graphics cards. The result marked a turning point. That combination, massive amount of data, neural networks, and GPU computing power became the golden recipe for modern AI and cemented Li's legacy.
2010年,李飞飞将ImageNet变成了一场竞赛。谁能构建出最好的算法来正确识别最多的图像?2012年,杰弗里·辛顿的多伦多大学团队携AlexNet参赛,这是一个由NVIDIA显卡驱动的算法。结果标志着一个转折点。海量数据、神经网络和GPU计算能力的结合,成为现代AI的黄金配方,并奠定了李飞飞的传奇地位。
What do you feel when you look at this? It's not the eight of us. It's the entire humanity. That's really the main story here. Because this comes from the photo of, what, about a hundred years ago when industrialization was in the booming phase and big urban skylines are being built. That was a civilizational moment. It's amazing that you're being recognized here, but like, could they have centered the photo? Come on.
你看到这张照片时有什么感受?这不是我们八个人,而是全人类。这才是这里真正的主线。因为这张照片来自大约一百年前,当时工业化正处于蓬勃发展阶段,大型城市天际线正在拔地而起。那是一个文明时刻。你能在这里被认可,这太神奇了,但是,他们就不能把照片居中吗?拜托。
It's true because there's more space here. I asked you about being the godmother of AI a few years ago. You did. How do you feel about that title? You know, Emily, I would never call myself godmother of anything. Remember, Emily, when you asked me, I was like, oh. I was taken aback because I don't naturally think about myself as godmother of anything. That's not my personality. I focus on work. If I rejected that on that spot, we once again would be in a situation where women just don't get recognized the same way as men.
确实,因为这里还有更多空间。几年前我问过你关于被称为“AI教母”的感受。你确实被这样称呼了。你对这个头衔有什么感觉?你知道吗,艾米丽,我永远不会称自己为任何事物的教母。记得吗,艾米丽,当你问我时,我当时是,哦。我吃了一惊,因为我自然不会把自己想成任何事物的教母。那不是我的性格。我专注于工作。如果我当场拒绝,我们就会再次陷入女性得不到与男性同等认可的局面。
I want more women to be called godmother of whatever creation, innovation, discovery that they have done. So do you embrace the term now or? Embrace is a big word. I don't go around and say, "Hi, I'm Fei-Fei I'm the godmother of AI." But people have adopted that term thanks to you. Artificial intelligence had long been discussed in policy and tech circles as something perpetually right around the corner. Until ChatGPT, AI went from a niche research field to the most talked about technology on the planet.
我希望更多女性能因她们的创造、创新、发现而被称作教母。那么你现在接受这个称呼了吗?接受是个很重的词。我不会到处说:“嗨,我是李飞飞,我是AI教母。”但人们已经采用了这个称呼,这要感谢你。人工智能在政策和科技圈中长期以来一直被讨论为总是即将到来的东西。直到ChatGPT出现,AI从一个冷门研究领域变成了地球上最受关注的技术。
Large language models, image generators, video generators. Suddenly it seemed like everyone was or wanted to be in the AI business. And let's be honest, things have come pretty far. Li saw an opportunity and launched her own startup in 2024. What's it like running your own shop? So here's the thing. I'm by far the most senior in the company. From my co-founders to my engineers and researchers and scientists, they're so talented, but they're also young.
大型语言模型、图像生成器、视频生成器。突然间,似乎每个人都在或想进入AI行业。说实话,事情已经发展得相当远了。李飞飞看到了机会,并于2024年创办了自己的初创公司。经营自己的公司是什么感觉?事情是这样的。我是公司里资历最深的。从我的联合创始人到我的工程师、研究员和科学家,他们都非常有才华,但他们也很年轻。
I do feel like a mom. And oftentimes I feel like a tiger mom. I do run this shop as, I would say, high standard. Li is working on advancing world models. If ImageNet taught AI how to see, this technology aims to guide AI through the physical world because describing how to catch a ball and actually catching one are very different things. And super intelligence, she's betting that lies beyond chatbots. The rest of the AI industry is focusing on LLMs.
我确实感觉自己像个妈妈。而且很多时候,我感觉自己像个虎妈。我确实以高标准来经营这家店,我会这么说。李飞飞正在推进世界模型的研究。如果ImageNet教会了AI如何看,这项技术旨在引导AI穿越物理世界,因为描述如何接住一个球和真正接住一个球是完全不同的事情。而超级智能,她押注它超越了聊天机器人。AI行业的其他部分正聚焦于大语言模型(LLMs)。
What can't LLMs do? If we want to create a world where we continue to push scientific discovery, where we want machines like robots to be a partner to us, all this language cannot do it alone. Building world model and spatial intelligence is not about anti - LLM. It's about the next frontier and the next chapter. So what exactly is a world model? World model at this point is an overloaded term. We look at spatial intelligence as three different kind of functions.
大语言模型(LLMs)不能做什么?如果我们想创造一个世界,在那里我们继续推动科学发现,在那里我们想让机器人这样的机器成为我们的伙伴,所有这些语言都无法单独做到。构建世界模型和空间智能并不是反对LLM。这是关于下一个前沿和下一篇章。那么,世界模型到底是什么?世界模型在这一点上是一个被过度使用的术语。我们将空间智能视为三种不同的功能。
One is rendering, one is simulation, and the third one is planning. Rendering is the world model outputs beautiful pixels for humans to consume. Most likely you're thinking about [OpenAI's] Sora. The second class of world model is they do simulation, but it doesn't serve just humans. It serves machines. They are focused on capturing the actual structure of the world, the geometry structure based on physics. The third kind
一种是渲染,一种是模拟,第三种是规划。渲染是世界模型输出美丽的像素供人类消费。很可能你想到的是[OpenAI的]Sora。第二类世界模型是做模拟,但它不仅仅服务于人类。它服务于机器。它们专注于捕捉世界的实际结构,基于物理的几何结构。第三种
原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力