Enfold: Embeds World Models into Representations, Slashing Robot Control Latency by 10x
Weili Zeng · hf · 2026-08-11
Current visual world generative models for robot control typically require executing a costly generative branch to predict the future. The Enfold framework proposes internalizing this future-generative computation into a representation predicted solely from current visual context and language instructions.
During training, multi-level intermediate states exposed by the generator supervise the encoder. At deployment, action prediction bypasses the generator entirely. Evaluated on LIBERO, RoboTwin2.0, and real-robot tasks, Enfold maintains strong control while reducing action latency by 3.7x compared to Fast-WAM (with Enfold-Flash achieving 10.1x). The representation effectively suppresses nuisance variations and adapts dynamically to human interventions, proving it goes beyond fixed trajectory replay.
More from Embodied
- Paper: DB-VIO, A Dual-Branch Framework for Visual Inertial Odometry — rsasaki0109 · 2026-08-11
- Embodied AI Startups Act as VCs: 29 Firms Make 125 Investments — FinanceYF5 · 2026-08-11
- Galaxy Z Fold 8 Hands-On: Light as a Passport, One-Hand Friendly, Makes Apple Feel Stuck in Past — bilawalsidhu · 2026-08-11
- Mistral Ventures into Robotics: Introduces VLM Navigation Research, Hits SOTA on R2RCE — sivareddyg · 2026-08-11
- Samsung Official Details Engineering Secrets of Galaxy Z Fold2 Hinge — MarwaEldiwiny · 2026-08-11
- INTACT by ZJU & Tsinghua: Robots Skip Trial-and-Error to Act Directly on Intent — jiqizhixin · 2026-08-11