The Causal Architecture of LingBot-VA 2.0
rohanpaul_ai · x · 2026-07-14
This post breaks down the core architecture of LingBot-VA 2.0:
- The control policy is trained in a causal manner from the ground up, rather than retrofitting a bidirectional video model into a controller.
- The video stream utilizes 128 routed experts, top-8 routing, and 1 shared expert. The total parameter count is around 13B, but only 1.9B are activated per token.
- The action stream remains dense, aiming to balance overall capability with control latency.
Related event: LingBot-VA 2.0: A Control-Native Foundation Model for Robotics(9 posts)→
More from Embodied
- Tesla expands Robotaxi rides to seven areas, including new Orlando and Tampa zones — elonmusk · 2026-07-22
- Hands-on robotics workshop on Saturday may be the last in-person session before August — StewartalsopIII · 2026-07-22
- NVIDIA pitches World Foundation Models as a way to scale physical AI data generation — MonaJalal_ · 2026-07-22
- RoboMME Podcast Preview: Benchmarking Memory for Robotic Policies — chris_j_paxton · 2026-07-21
- Gritt says an 8-person crew now installs 3,000 to 4,000 solar panels a day — HaktanSuren · 2026-07-21
- A helium-powered flying robot whale aims to be a quiet companion pet — chris_j_paxton · 2026-07-21