CD-LAM Framework Boosts Robot Video Learning Efficiency 12x

jiqizhixin · x · 2026-08-08

Aether AI and UC San Diego introduced CD-LAM (Causally Debiased Latent Action Model), a new training framework to fix action confusion when AI models learn robot skills from video.

Existing models often conflate core actions with irrelevant backgrounds or clutter. CD-LAM tackles this by applying three debiasing tricks during fine-tuning, forcing the model to focus solely on action-relevant dynamics to extract cleaner, controllable latent actions.

The framework outperforms the DreamDojo baseline in action following, visual fidelity, and efficiency. At the 14B scale, it matches DreamDojo using over 12 times fewer robot action adaptation updates, and also delivers strong gains at the 2B scale.

Original post →

More from Embodied

Embodied channel →