Google如何教LLM控制机器人并引发行业热潮
How Google taught LLMs to control robots and started a robotics boom
深度复盘具身智能起源与产业演变,理清了RT-2的技术逻辑及团队创业脉络,对理解当前机器人赛道格局极具参考价值。
It’s day three of Robot Week! You can click here to get 25% off an annual subscription.
今天是机器人周第三天!点击此处可享年度订阅八五折优惠。
Most people first heard about large language models after OpenAI introduced ChatGPT in 2022. But in the minds of many AI researchers, the key breakthrough came two years earlier with the release of GPT-3.
大多数人是在2022年OpenAI推出ChatGPT之后才首次听说大语言模型的。但在许多AI研究人员看来,真正的关键突破早在两年前随着GPT-3的发布就出现了。
With 175 billion parameters, the OpenAI model was more than 100 times larger than its predecessor, GPT-2. It was trained on a massive 300 billion tokens. And as a result, it generalized far better than previous models. For the first time, a single model could perform a wide variety of tasks — from translating between languages to answering trivia questions — without task-specific training.
拥有1750亿参数的OpenAI模型比其前身GPT-2大了100多倍。它在海量的3000亿个token上进行了训练。因此,它的泛化能力远优于之前的模型。这是第一次,单个模型无需针对特定任务进行训练,就能执行各种任务——从跨语言翻译到回答常识性问题。
It took OpenAI two more years to develop the techniques that transformed this raw “base model” into a user-friendly chatbot like ChatGPT. And then it took a couple more years to develop the techniques — like long-context reasoning, tool use, and context management — that transformed those early chatbots into the powerful agents we have today.
OpenAI又花了两年时间开发技术,将这个原始的“基础模型”转变为用户友好的聊天机器人,如ChatGPT。随后又花了几年时间开发长上下文推理、工具使用和上下文管理等技术,将这些早期的聊天机器人转变为今天我们所拥有的强大智能体。
In short, there was a long road from GPT-3 in 2020 to Claude Code in 2025. But for those who knew where to look, the potential of LLMs was already clear in 2020.
简而言之,从2020年的GPT-3到2025年的Claude Code,这条路很长。但对于知道该看哪里的人来说,LLM的潜力在2020年就已经很清晰了。
The robotics world is now traveling a similar path. Its “GPT-3 moment” came in July 2023, when Google announced a model called RT-2. To create it, Google trained a multimodal LLM to directly generate robot actions. RT-2 wasn’t Google’s first transformer-based robotics model — the company released a predecessor called RT-1 a few months earlier, for example — but RT-2 was massively larger than earlier models. RT-1 had 35 million parameters. The RT-2 models had billions of parameters.
如今,机器人领域正在走一条相似的道路。它的“GPT-3时刻”出现在2023年7月,当时Google宣布了一个名为RT-2的模型。为了创建它,Google训练了一个多模态大语言模型来直接生成机器人动作。RT-2并不是Google首个基于Transformer的机器人模型——例如,该公司几个月前发布了其前身RT-1——但RT-2比早期模型庞大得多。RT-1有3500万个参数,而RT-2模型则有数十亿个参数。
And as with GPT-3, size mattered. The RT-2 team reported its model showed “significant improvements to generalization over objects, scenes, and instructions.” They added that the new model exhibited “a breadth of emergent capabilities inherited from web-scale vision-language pretraining.”
与GPT-3一样,规模至关重要。RT-2团队报告称,其模型在物体、场景和指令的泛化方面显示出“显著改进”。他们补充说,新模型展现了“源自网页级视觉-语言预训练的广泛涌现能力”。
For example, researchers placed a can of Coca-Cola on a counter alongside framed photos of Snoop Dogg, Tom Cruise, and Taylor Swift. They then prompted the robot to “move coke can to Taylor Swift.” The robot grabbed the can and moved it toward Swift’s photo.
例如,研究人员将一罐可口可乐放在柜台上,旁边是Snoop Dogg、Tom Cruise和Taylor Swift的带框照片。然后他们提示机器人“把可乐罐移到Taylor Swift那里”。机器人抓起可乐罐并移向Swift的照片。
At the time, Karol Hausman was a member of the RT-2 team. In a March interview, he described this as a moment of “huge, huge excitement” because “the robot models had never had any of Taylor Swift in their data. It had to understand the concept of Taylor Swift, connect it to the image of Taylor Swift, and then connect it to the right motion that would move the Coke can to the picture of Taylor Swift, all from Internet data.”
当时,卡罗尔·豪斯曼(Karol Hausman)是 RT-2 团队的成员。他在三月的一次采访中称,这是一个“巨大、巨大的兴奋”时刻,因为“机器人模型的数据中从未包含过任何泰勒·斯威夫特(Taylor Swift)的内容。它必须理解泰勒·斯威夫特的概念,将其与泰勒·斯威夫特的图像联系起来,然后将其与正确的动作联系起来,从而将可乐罐移动到泰勒·斯威夫特的图片上,这一切都源自互联网数据。”
“That was the moment where it clicked for us that it could actually work — where you could bring in a lot of prior knowledge from LLMs, from the Internet, and connect it to robot motions,” Hausman said.
豪斯曼表示:“那是我们意识到它真正可行的时刻——你可以从大型语言模型(LLMs)和互联网中引入大量先验知识,并将其与机器人动作联系起来。”
Google dubbed RT-2 a vision-language-action (VLA) model. Both Google’s approach and the term VLA quickly became industry standards. But as impressive as RT-2 was, it also had significant shortcomings — shortcomings the industry has been working to remedy over the last three years.
谷歌将 RT-2 称为视觉-语言-动作(VLA)模型。谷歌的方法以及 VLA 这一术语迅速成为行业标准。但尽管 RT-2 令人印象深刻,它也存在重大缺陷——这些缺陷正是行业在过去三年中致力于弥补的。
The RT-2 breakthrough kicked off a robotics boom that’s been underway ever since. Big companies in both the US and China have poured resources into robotics. Numerous robot startups have been created, and several have raised hundreds of millions of dollars in venture capital. And the models powering most of these robots are based on the basic architecture Google pioneered back in 2023.
RT-2 的突破引发了一场至今仍在持续的机器人热潮。美国和中国的各大公司纷纷向机器人领域投入资源。众多机器人初创企业应运而生,其中几家已筹集到数亿美元的创业投资。而驱动大多数这些机器人的模型,均基于谷歌在 2023 年开创的基础架构。
The origins of RT-2, the first VLA model
首个 VLA 模型 RT-2 的起源
The robot Google used to train RT-2. (Image courtesy of Google)
谷歌用于训练 RT-2 的机器人。(图片来源:谷歌提供)
Google invented the transformer in 2017 and had been experimenting with large language models ever since. The company had also been working on robotics for many years. So combining LLMs and robots was an obvious research direction.
谷歌于 2017 年发明了 Transformer,并自此一直在探索大型语言模型。该公司多年来也一直在从事机器人研究。因此,将 LLMs 与机器人结合是一个显而易见的研究方向。
In March 2023, Google announced PaLM-E, a 12-billion-parameter model that was optimized for robotics (the “E” stood for “embodied”). PaLM-E was a vision-language model (VLM) — meaning an LLM trained to understand images as well as text. It had been trained to generate natural-language robot commands like “move the blue block to the left.”
2023 年 3 月,谷歌发布了 PaLM-E,这是一个专为机器人优化的 120 亿参数模型(“E”代表“具身”,embodied)。PaLM-E 是一个视觉-语言模型(VLM),即一个经过训练以同时理解图像和文本的 LLM。它被训练生成自然语言的机器人指令,例如“将蓝色方块向左移动”。
But PaLM-E couldn’t control a robot directly. Google’s robots didn’t have enough onboard computing power to run a VLM as large as PaLM-E. So PaLM-E ran in the cloud, and it was designed to work with a second, smaller model that would run on the robot. This second model would translate PaLM-E’s English instructions into low-level robot commands.
但 PaLM-E 无法直接控制机器人。谷歌的机器人没有足够的机载计算能力来运行像 PaLM-E 这样庞大的 VLM。因此,PaLM-E 在云端运行,并设计为与第二个较小的模型协同工作,后者将在机器人上运行。这个第二模型会将 PaLM-E 的英语指令转换为底层机器人命令。
The RT-2 team’s plan was simple: delete the smaller model and instead train PaLM-E to directly control the robot.1 RT-2 — like PaLM-E — was too big to run directly on a robot. So the team ran the model in a Google data center and had it send commands to the robot over the network.
RT-2 团队的计划很简单:删除较小的模型,转而训练 PaLM-E 直接控制机器人。1 RT-2 —— 与 PaLM-E 一样 —— 体积过大,无法直接在机器人上运行。因此,团队在 Google 数据中心运行该模型,并通过网络向机器人发送指令。
Like any LLM, RT-2 worked by prompting. Google would send RT-2 a prompt like “What action should the robot take to move coke can to Taylor Swift?” along with an image from the robot’s camera.
像任何大语言模型(LLM)一样,RT-2 通过提示词工作。Google 会向 RT-2 发送类似“机器人应采取什么行动将可乐罐移到泰勒·斯威夫特那里?”的提示词,并附带来自机器人摄像头的图像。
RT-2 would respond with a sequence of numbers like “1 128 91 241 5 101 127 217.” The robot would interpret this as a command to move the robot’s gripper to certain x-y-z coordinates (like x=128, y=91, and z=241), rotate the gripper to a certain angle (roll=5, yaw=101, pitch=127), and open (or close) the gripper to a certain position (217).
RT-2 会响应一串数字,例如“1 128 91 241 5 101 127 217”。机器人会将此解释为指令,将机械臂夹爪移动到特定的 x-y-z 坐标(如 x=128、y=91、z=241),将夹爪旋转到特定角度(滚转=5、偏航=101、俯仰=127),并将夹爪打开(或关闭)到特定位置(217)。
Then RT-2 would get the same prompt again, but with a fresh image. The model would generate another sequence of numbers representing a new target position for the robot arm. The robot would move its arm another few inches. Then the whole cycle would repeat again. It might take dozens of iterations to complete a task like “move coke can to Taylor Swift.”
然后 RT-2 会再次收到相同的提示词,但附带一张新的图像。模型会生成另一串数字,代表机械臂的新目标位置。机器人会将手臂再移动几英寸。随后整个循环重复进行。完成诸如“将可乐罐移到泰勒·斯威夫特那里”这样的任务可能需要数十次迭代。
To transform PaLM-E into RT-2, Google had to teach the model how to generate low-level robot instructions. That required a different kind of training data.
为了将 PaLM-E 转化为 RT-2,Google 必须教会模型如何生成底层机器人指令。这需要不同类型的训练数据。
To collect that data, Google built three test kitchens and purchased 13 robots. Over the course of 17 months, human workers teleoperated the robots as they performed tasks — picking up objects, opening drawers, placing objects in the drawers, and so forth — more than 130,000 times.
为了收集这些数据,Google 建造了三间测试厨房并购买了 13 台机器人。在 17 个月的时间里,人类工作人员远程操控机器人执行任务——拿起物体、打开抽屉、将物体放入抽屉等——超过 130,000 次。
Training PaLM-E on this data gave RT-2 surprisingly broad capabilities. Robots could manipulate objects they hadn’t seen before. They could operate in new kitchens. And they could complete tasks on counters that were cluttered with “distractor objects” that weren’t needed for the assigned task.
用这些数据训练 PaLM-E 赋予了 RT-2 令人惊讶的广泛能力。机器人可以操作它们从未见过的物体。它们可以在新的厨房中操作。并且它们可以在杂乱无章的台面上完成任务,这些台面上布满了与指定任务无关的“干扰物体”。
Five roboticists left Google to co-found Physical Intelligence
五名机器人学家离开 Google,共同创立了 Physical Intelligence
Karol Hausman was excited by the RT-2 breakthrough, but he also concluded that Google wasn’t the right place to develop the technology.
Karol Hausman 对 RT-2 的突破感到兴奋,但他同时也得出结论,认为 Google 不是开发这项技术的合适场所。
“It became clear that the way to accomplish this is to create an organization whose sole purpose is to solve physical intelligence,” Hausman said in March. “It can’t be solved as priority number 20 in another organization.”
Hausman 在三月份表示:“很明显,实现这一目标的方法是创建一个唯一目的就是解决物理智能的组织。它不能作为另一个组织中的第 20 优先级来解决。”
So Hausman became the CEO of a startup called Physical Intelligence. He was joined by four other members of Google’s RT-2 team and two others from outside Google.
因此,Hausman 成为了一家名为 Physical Intelligence 的初创公司的首席执行官。他与 Google RT-2 团队的其他四名成员以及两名来自 Google 外部的人士共同加入了该公司。
According to Hausman, the team sought out “investors that are fully aligned with this starting as a research company and not being oriented around short-term revenue.”
据豪斯曼(Hausman)介绍,团队寻找的是“那些完全认同公司起步阶段作为研究机构、不以短期收入为导向的投资者。”
“If we do this right, this is going to completely change the world and it’s going to be the most valuable business of all time,” Hausman said. “But you need to have the patience to let us do it the right way.”
豪斯曼表示:“如果我们做对了,这将彻底改变世界,并成为有史以来最有价值的企业。但你需要有耐心,让我们以正确的方式去做。”
There was a lot to do. RT-2 was a big improvement over previous robotic models, but it was still far less capable than the average human. Over the last two years, the Physical Intelligence (PI) team has been working hard to close that gap. The company has been remarkably transparent, publishing at least 10 papers describing their work. For this story, I read all the PI papers I could find — along with 20 more from other companies and academic labs.
要做的事情还有很多。RT-2 相比之前的机器人模型有了显著提升,但其能力仍远逊于普通人。在过去两年里,Physical Intelligence (PI) 团队一直在努力缩小这一差距。该公司一直保持着极高的透明度,发表了至少 10 篇描述其工作的论文。为了撰写这篇报道,我阅读了所有能找到的 PI 论文,以及来自其他公司和学术实验室的另外 20 篇论文。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力