LeCun's H-JEPA: Hierarchical world model lifts planning success from 18% to 73%
randall_balestr · x · 2026-10-06
H-JEPA is the first end-to-end trained hierarchical world model for long-horizon visual planning, from a team including Yann LeCun and Randall Balestriero (arXiv:2610.06805).
- Architecture: a hierarchy of action-conditioned JEPAs, each level predicting farther ahead in its own learned latent space; planning proceeds top-down, with each level's predictions serving as subgoals for the level below.
- Key mechanism: SIGReg/LeWM prevents dimensional collapse and keeps latent spaces well-conditioned, so semantic abstractions emerge naturally—slow, task-relevant state is retained while fast unpredictable detail is discarded.
- Results: across four simulated navigation/manipulation environments, hierarchical planning beats flat JEPA; on Visual AntMaze, a three-level hierarchy raises success from 18% to 73% with less planner compute. Ablations attribute gains to both temporal decomposition and higher-level goal representations.
- Real robots: with inverse-dynamics supervision, the approach extends to DROID real-robot videos, improving offline planning fidelity at lower compute.
Paper and code are open-sourced.
Related event: LeCun Team Releases H-JEPA, an End-to-End Hierarchical World Model(5 posts)→
More from Embodied
- Minerva, a robot startup for extreme jobs, comes out of stealth — kyliebytes · 2026-10-07
- Rhoda's Single Robot Model Handles Fridge and Washing Machine Tasks — GordonWetzstein · 2026-10-07
- Robot demo shows gentle start at household chores — Sentdex · 2026-10-07
- ProgressCompass: context injection cuts embodied progress reward model error by 63% — Jianshu Zhang · 2026-10-07
- Video2Skill benchmark: most of 19 open-source VLMs fail at streaming embodied skill discovery — Jianshu Zhang · 2026-10-07
- MM-ABC robot foundation model hits 83% on real-world mobile manipulation tasks — Qiwei Liang · 2026-10-07