跳到主内容
精选80Rohan Paul模型发布/更新多源精选 ×2

LingBot-VLA 2.0:跨20种机器人构型的全身控制策略

Most VLAs (Vision-Language-Action Models) handle task variety only inside narrow…

原文

Most VLAs (Vision-Language-Action Models) handle task variety only inside narrow, fixed bodies;

LingBot-VLA 2.0 from @robbyant_brain trains one policy across 20 configurations with whole-body control.

Also to avoid the damaging noise of a robot datasets, LingBot-VLA 2.0 filters 90,000 raw hours into 50,000 cleaner high-quality real-robot data hours.

Supports 20 robot configurations and whole-body degree-of-freedom control.

So LingBot-VLA 2.0 from @robbyant_brain is a whole-body robot policy. It controls arms, grippers, dexterous hands, the head, waist, and mobile base through one 55-dimensional action format, so the same model can learn from very different robot bodies.

About 90,000 raw robot hours are checked for jerky motion, broken signals, camera mismatch, blur, dropped frames, and long static periods before 50,000 hours remain.

Human videos are filtered for hand-object interaction, then camera motion and hand pose are reconstructed as action data.

A sparse Mixture of Experts module lets each action token use a few specialized networks while a shared expert preserves common skills.

Another training signal asks the model to predict current and future depth plus video features, pushing it to track object geometry and likely scene changes during the next action chunk.

On Agilex GM-100, it reaches 66.2% progress and 34.4% success, ahead of pi0.5 at 59.1% and 32.2%. It also leads pi0.5 across both long-horizon mobile tasks in both in-domain and out-of-distribution tests, although several individual GM-100 tasks still favor other models.

🧵 1.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近