跳到主内容
精选85Rohan Paul模型发布/更新多源精选 ×4

首个开源具身场景视频生成模型LingBot-Video发布

A first-of-its-kind open-source video generation model for embodied scenarios, u…

原文
推荐理由

做具身智能或视频生成的同学注意了,这是首个开源且专为具身场景设计的视频生成模型,物理合理性评估方法值得细读,建议跑一下 RBench 对比效果。

A first-of-its-kind open-source video generation model for embodied scenarios, using a Mixture-of-Experts design.

LingBot-Video from @robbyant_brain, a Diffusion Transformer video foundation model built for embodied AI.

  • The flagship 30B model activates only 3B parameters per generation, at 1M-token sequences
  • At the same total parameter scale it reaches ~3X the inference efficiency of a Dense architecture at 1M-token sequences (3.18× vs. Dense 30B)
  • Aesthetic-only reward models are why video generators fake physics, and this paper basically says so. Physical plausibility is not treated as vibes here; it gets scored through causality, object permanence, non-penetration and material realism.
  • The model was not trained only on regular internet videos. Its pretraining data also includes 70,000+ hours of embodied footage, covering robot manipulation, navigation, and first-person action videos, so the model sees examples where movement actually changes the world.
  • On RBench, it reports a 0.620 average, ahead of every listed open and closed model, while still losing some columns to Wan 2.6.

🧵 1.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近