Ant Group Proposes LingBot-VA 2.0 Video-Action Model
jiqizhixin · x · 2026-07-18
Ant Group introduced LingBot-VA 2.0, a video-action foundation model built from scratch specifically for real-world control.
Unlike standard video models designed for digital content, this model introduces four breakthroughs:
- Semantic-Action Tokenizer: Links the robot's visual input with the actions it should execute.
- Causal Training Method: Prevents the model from forgetting previously executed actions.
- Sparse MoE Backbone: Increases model capacity without compromising operational speed.
- Asynchronous Inference Scheme: Predicts future steps in real-time while executing the current action.
Results: Its few-shot generalization capabilities in complex manipulation tasks surpass previous video-action models.
More from Embodied
- Humanoid robots are moving from labs into public culture — Olivier__OG · 2026-07-21
- Polymarket puts Tesla’s California robotaxi launch odds at 16% this year — Polymarket · 2026-07-21
- Tesla expands robotaxi service to Orlando and Tampa — Polymarket · 2026-07-21
- Humanoid robots are approaching a deeper uncanny valley — GlenBradley · 2026-07-21
- Polymarket gives Tesla’s Optimus just a 17% chance of debuting this year — Polymarket · 2026-07-21
- UK robotics startup Humanoid raises $152 million at a $1.35 billion valuation — Polymarket · 2026-07-21