Enfold: Embeds World Models into Representations, Slashing Robot Control Latency by 10x

Weili Zeng · hf · 2026-08-11

Current visual world generative models for robot control typically require executing a costly generative branch to predict the future. The Enfold framework proposes internalizing this future-generative computation into a representation predicted solely from current visual context and language instructions.

During training, multi-level intermediate states exposed by the generator supervise the encoder. At deployment, action prediction bypasses the generator entirely. Evaluated on LIBERO, RoboTwin2.0, and real-robot tasks, Enfold maintains strong control while reducing action latency by 3.7x compared to Fast-WAM (with Enfold-Flash achieving 10.1x). The representation effectively suppresses nuisance variations and adapts dynamically to human interventions, proving it goes beyond fixed trajectory replay.

Original post →

More from Embodied

Embodied channel →