印度成AI机器人数据收集中心,美企高薪招募劳工
How India Is Training the Robots of the Future | Big Take Asia
揭示了具身智能产业链上游的关键要素——高质量物理数据的获取路径与成本结构。对于关注AI硬件、机器人及算力基础设施的投资人,这提供了关于供应链转移和数据要素价值的实证参考。
Bloomberg Audio Studios, podcasts, radio, news. On top of a mountain of discarded plastic, a 43-year-old woman named Senita Ratau is hard at work. She's in a bright yellow sari with green glass bangles sifting through piles of litter, plastic waste, which she'll sort and clean for recycling. Senita is based in New Delhi, but scenes like this one take place every day all over India on construction sites, in homes and recycling plants where manual labor is done.
彭博音频工作室、播客、广播、新闻。在一座由废弃塑料堆积成的小山上,43岁的塞尼塔·拉图正在辛勤工作。她身穿亮黄色的纱丽,手腕上戴着绿色玻璃手镯,正从成堆的垃圾和塑料废物中分拣出可回收物进行清理。塞尼塔常驻新德里,但类似这样的场景每天都在印度各地的建筑工地、家庭和回收工厂里上演,那里全靠人工劳作。
Nothing out of the ordinary except for one thing.
一切看似平常,除了其中一件事。
Suddenly, you saw this high-tech contraption on her forehead, an iPhone.
突然,你看到她额头上戴着一个高科技装置——一部iPhone。
An iPhone strapped to her head recording every move she makes. Stha Ry covers tech in India for Bloomberg. She met Senita at a recycling plant earlier this year. Senita has spent years sorting and cleaning discarded plastic for about $200 a month. Usually the factories where she works are lively places.
一部绑在她头上的iPhone,记录着她做的每一个动作。斯哈·瑞(Stha Ry)为彭博社报道印度的科技动态。今年早些时候,她在一家回收厂见到了塞尼塔。多年来,塞尼塔一直从事分拣和清理废弃塑料的工作,月薪约200美元。通常,她工作的工厂都是热闹非凡的地方。
There's a lot of banter. There's a lot of conversation. There's a lot of silliness that goes on. in this. It was deadly quiet because all these women were not allowed to talk to each other
这里充满了玩笑、交谈和各种嬉闹。然而这一次却死一般寂静,因为这些女性被禁止互相交谈。
8 hours a day in total silence. But that's not all.
整整8小时保持绝对沉默。但这还不是全部。
She was hunched forward constantly because she was not loved to capture any other part of her body except her wrist. It's exactly what uh a humanoid robot camera would see if it looked down. That was the motion she was asked to capture. Senita doesn't know what the footage is for. All she knows is that this phone recording earns her extra money. Money she uses to send her four kids to school. But the recordings she and her fellow workers make are critical to the future of AI and robotics.
她不得不始终保持身体前倾的姿势,因为拍摄要求只捕捉她的手腕,不允许拍到身体的其他部位。如果一个人形机器人摄像头向下看,看到的景象正是如此。这正是她被要求摆出的动作。塞尼塔并不知道这些 footage 的用途。她只知道,通过这部手机录制视频能让她多赚些钱,她用这笔钱供四个孩子上学。但她和工友们录制的影像,对于人工智能和机器人的未来至关重要。
They'll be fed into AI systems to teach humanoid robots how to move and navigate the physical world like us, one frame at a time. The goal really is to capture the micro gestures, the really tiny hand movements that, you know, robots need to learn to become better at everyday tasks. The world has started to embrace chat bots and agentic AI systems like Claude or Chat GBT which rely on cloud computing or local hardware to run.
这些影像将被输入AI系统,用于教人形机器人如何像我们一样移动和在物理世界中导航,一次一帧。真正的目标是捕捉细微的手势和极其微小的手部动作,你知道的,机器人需要学习这些才能在日常任务中表现得更好。世界已经开始拥抱聊天机器人和代理式AI系统(如Claude或ChatGPT),它们依赖云计算或本地硬件来运行。
You give an AI agent a voice, text prompt or task, it gives you an answer. And for years, the robotics industry has tried to create humanoid devices. AI that moves in the physical world the same way that we do. When it's gone well, it's been dazzling.
你给AI代理一个语音指令、文本提示或任务,它会给你答案。多年来,机器人行业一直在努力创造能够在物理世界中像人类一样行动的人形设备。当进展顺利时,其表现令人眼花缭乱。
Okay. Oh,
好的。哦,
we have a new world record.
我们创造了新的世界纪录。
At the World Humanoid Robot Games in Beijing last month, essentially an Olympics for machines, a Chinese robot shattered the 100 meter sprint record. Its time 8.64 seconds. Nearly a full second faster than Usain Bolt, the fastest man on Earth. But step outside that sterile controlled track, and it's a completely different story. When robots come into the real world, there's invariably a blooper reel awaiting the texture of real life, the unevenness of perhaps the terrain, the the spatial intelligence that's required by for these robots.
在上月于北京举行的世界人形机器人运动会——这本质上是一场机器界的奥运会——上,一款中国机器人打破了百米短跑纪录。其成绩为8.64秒,比地球上跑得最快的人尤塞恩·博尔特(Usain Bolt)快了将近整整一秒。但一旦走出那个无菌且受控的跑道,情况就完全不同了。当机器人进入现实世界时,总有一系列失误镜头在等待着它们,以应对真实生活的质感、地形的不平整,以及这些机器人所需的空间智能。
All of these things, you know, make it really difficult for robots to be trained in the lab. They have to be trained in the real world with real world data, which is where, you know, all of these data collectors come in. Cityroup projects the market for human-like robots could reach $7 trillion by 2050. And to learn how to perform tasks like a human, robots need huge amounts of so-called egocentric data. Footage captured directly through the eyes of a person.
所有这些情况,你知道,使得机器人在实验室里训练变得非常困难。它们必须在现实世界中,使用现实世界的数据进行训练,而这正是所有这些数据收集者发挥作用的地方。Cityroup预测,到2050年,类人机器人的市场规模可能达到7万亿美元。为了学会像人类一样执行任务,机器人需要海量的所谓“第一人称视角”数据,即直接通过人的眼睛捕捉的画面。
China has been the biggest generator of that data, but it's keeping it within its borders. So the demand has shifted to the world's most populous country.
中国一直是这类数据的主要生产国,但它将数据保留在国内。因此,需求转向了世界上人口最多的国家。
So India really has become the epicenter of data collection in the past few months. Tens of thousands of workers like Sonita uh shoe factory workers, denim factory workers, construction workers, all kinds of people have been recruited to record these firsterson camera footage or what is called egocentric footage. It was a sort of unsettling moment for me when I talked to these workers because most of them were recording themselves into their own job extinction.
因此,在过去几个月里,印度确实成为了数据收集的中心。成千上万的工人,如索尼塔(Sonita)这样的鞋厂工人、牛仔布厂工人、建筑工人,各行各业的人都已被招募来录制这些第一人称视角摄像机画面,也就是所谓的“第一人称视角” footage。当我与这些工人交谈时,那是一种令人不安的时刻,因为大多数人正在录制自己走向职业消亡的过程。
Welcome to the big take Asia from Bloomberg News. I'm Rebecca Chung Wilkins in for one. Every week we take you inside some of the world's biggest and most powerful economies and the markets tycoons and businesses that drive this ever shifting region. Today on the show, why the road to the next AI breakthrough runs through India and what this race for human data means for the US China tech rivalry. Humanoid robots and online chat bots both use AI.
欢迎来到彭博新闻社的《亚洲大趋势》节目。我是瑞贝卡·钟·威尔金斯(Rebecca Chung Wilkins),本期由我代班主持。每周我们带你深入全球最大、最具影响力的经济体内部,以及推动这个不断变化的地区的市场、大亨和企业。本期节目中,我们将探讨为何通往下一次AI突破的道路经过印度,以及这场争夺人类数据的竞赛对美中科技竞争意味着什么。人形机器人和在线聊天机器人均使用人工智能。
But Bloomberg Sera Ry says teaching a large language model to write an essay is completely different from teaching machine to cook a meal.
但彭博社的Sera Ry表示,教大型语言模型写文章与教机器做饭完全是两回事。
Real world humanoid robotics is a completely different ballgame. This requires real world data, actual human motion and human movements that have to be learned by the models to teach the robots. There is no availability of any of this data online, you know, on the internet.
现实世界的人形机器人是一个完全不同的游戏。这需要现实世界的数据,实际的人类动作和运动,这些必须由模型学习,以便教会机器人。在互联网上,你根本找不到任何此类数据的可用性。
You might think, wait, the internet has plenty of data. All those videos on YouTube or Tik Tok, all those humans doing random things. But
你可能会想,等等,互联网上有很多数据。YouTube 或 TikTok 上的所有视频,以及所有那些正在做随机事情的人类。但是
it's really raw data and it's not firsterson egocentric data. It is not the data that a humanoid robot's camera would see when it's looking down or looking forward. That kind of data has to be collected in a highly standardized way. Faces can't appear in the footage and conversations aren't allowed. Equally important, the data has to come from as many different humans as possible.
它确实是原始数据,但不是第一人称自我中心(egocentric)数据。这不是人形机器人摄像头在向下看或向前看时会看到的数据。这类数据必须以高度标准化的方式收集。画面中不能出现人脸,也不允许有对话。同样重要的是,数据必须尽可能来自不同的人类。
Each person has their own way of doing things. Like how you perhaps open a can be completely different from how I open a can. So each person's data is valuable and multiple labs are looking for what is called proprietary data.
每个人都有自己做事的方式。比如你开罐头的方式可能和我完全不同。因此,每个人的数据都是有价值的,多个实验室都在寻找所谓的专有数据。
Sura says India is an ideal place to gather this kind of data. The country has an enormous population. Labor is cheap and workers still do a lot of manufacturing by hand. the diversity whether it's the shoe factories, the denim uh factories, whether it's the you know warehouses of different kinds, whether it's the households, fishing villages or the farms or construction sites or the mines. It's incredible the diversity of tasks that can be recorded in India.
Sura 表示印度是收集此类数据的理想之地。该国人口庞大。劳动力便宜,工人仍然大量手工进行制造。无论是鞋厂、牛仔布工厂,还是各种类型的仓库,亦或是家庭、渔村、农场、建筑工地或矿山,其多样性令人难以置信。在印度可以记录的任务多样性也是惊人的。
No other country could offer this kind of a diversity of manual labor. Not even China I would say because a lot of these tasks particularly in factories and warehouses are already automated in some way or the other in China.
没有其他国家能提供如此多样的人工劳动。我甚至会说连中国也不行,因为在中国,许多这些任务,特别是在工厂和仓库中的任务,已经在某种程度上实现了自动化。
And what about the cost? Can you break down what kind of cost you you have if you're buying this data in India versus somewhere else say Singapore or even the US?
那么成本呢?你能分解一下,如果在印度购买这些数据,与在新加坡甚至美国等地相比,成本构成是怎样的吗?
I haven't looked at uh the costs in Singapore but I can definitely talk about the costs uh in India versus the US. For instance, the raw first person footage of the egocentric data that India provides could come for about $7 an hour. Uh but if it's cleaned and annotated, it could go for as much as $30 an hour. That's a huge value addition right there. And the same data, you know, could go from two to three to five times more expensive in the US when it's collected in the US.
我没有研究过新加坡的成本,但我肯定可以谈谈印度与美国之间的成本差异。例如,印度提供的自我中心数据的原始第一人称片段,每小时成本约为 7 美元。但如果经过清洗和标注,价格可能高达每小时 30 美元。这已经是一个巨大的附加值。而且,同样的数据,在美国收集时,价格可能是印度的两到三倍,甚至五倍。
Tesla, for instance, pays $48 an hour in Palo Alto for um motion capture data collector inhouse.
例如,特斯拉在帕洛阿尔托为内部动作捕捉数据收集员支付每小时 48 美元。
The factory floor is just the starting point for these egocentric videos. Once the footage is captured, it's sent to local firms that clean it up, check the quality, and label the data so it can be used to train AI models. Demand for this work has surged, helping fuel a boom in India's AI services sector with a growing number of companies built around processing and supplying data to tech firms. Who are these startups working with?
工厂车间只是这些自我中心式视频的起点。一旦拍摄完成,素材就会被发送给当地公司进行清理、质量检查和数据标注,以便用于训练人工智能模型。此类工作的需求激增,推动了印度人工智能服务行业的繁荣,越来越多的公司围绕为科技公司处理和提供数据而建立。这些初创公司与哪些企业合作?
What's the sort of end point of all of this proprietary data?
所有这些专有数据的最终用途是什么?
Now, none of these companies were able to tell me who exactly was the end customer. Uh not that they didn't know, but the NDAs were so strict and so binding that none of them were allowed to reveal. But enough hints were dropped that these were all like major American and global companies including Google and Tesla and uh Nvidia and Meta. You know, Nvidia for instance is now getting into uh physical AI and embodied
目前,这些公司都无法告诉我具体的终端客户是谁。并非他们不知道,而是保密协议(NDA)极其严格且具有约束力,不允许他们透露。但足够的暗示表明,这些都是主要的美国及全球性公司,包括谷歌、特斯拉、英伟达和Meta。例如,英伟达现在正涉足物理人工智能和具身智能领域。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力