D-JEPA proposes decision-aligned latent world model, hits 87.89% on PushT

NEBULIS-Lab · hf · 2026-09-29

NEBULIS-Lab introduces D-JEPA, a decision-aligned latent world model.

Problem: latent world models predict action consequences, but accurate prediction doesn't guarantee latent distance reflects which candidate will execute successfully — the paper identifies a "decision-local prediction gap" where candidates predicted closer to the goal can yield worse realized outcomes.

Method: D-JEPA learns decision-relevant relations among candidate futures from executed outcomes, using a bounded permutation-equivariant operator that jointly reasons over goal-relative predictive features and ordinal evidence, refining pretrained predictive geometry where action choices matter most. It realizes this structure in JEPA-compatible representations for native latent-distance planning.

Results: gains across latent control, manipulation, pretrained action models, physical robots and autonomous driving — 87.89% success on PushT, +15.04 avg on RoboTwin, +17 points on physical robot tasks.

Original post →

More from Embodied

Embodied channel →