CD-LAM Framework Boosts Robot Video Learning Efficiency 12x
jiqizhixin · x · 2026-08-08
Aether AI and UC San Diego introduced CD-LAM (Causally Debiased Latent Action Model), a new training framework to fix action confusion when AI models learn robot skills from video.
Existing models often conflate core actions with irrelevant backgrounds or clutter. CD-LAM tackles this by applying three debiasing tricks during fine-tuning, forcing the model to focus solely on action-relevant dynamics to extract cleaner, controllable latent actions.
The framework outperforms the DreamDojo baseline in action following, visual fidelity, and efficiency. At the 14B scale, it matches DreamDojo using over 12 times fewer robot action adaptation updates, and also delivers strong gains at the 2B scale.
More from Embodied
- Dyna Robotics Teases 'Most Exciting' Robotics Breakthrough, Not Just a Demo — JasonMa2020 · 2026-08-08
- CoRL 2026 Workshop on Modeling Uncertainty in Robotic World Models Announced — mengyer · 2026-08-08
- Wayve Unveils GAIA-4 World Model: Solving Safety Simulation for Autonomous Driving — alexgkendall · 2026-08-08
- Google shows off Gemini Robotics ER 2 in new demo — ColbyHawker · 2026-08-08
- Reward Shaping Accelerates Robot RL Training by 5x — carlosdponx · 2026-08-08
- Partial Autonomy is the Most Reliable Strategy for Robotics Startups — ishabytes · 2026-08-08