PhyLatent: Optimizing JEPA World Model Representations for Better Robot Control

burny_tech · x · 2026-08-10

To address the shortcomings of Joint-Embedding Predictive Architecture (JEPA) world models in control tasks, the paper PhyLatent introduces a novel dynamics-relevant training objective.

The author notes that merely preventing global latent collapse is insufficient to ensure the model accurately captures physical states and action consequences. The researchers identified three failure modes: physical invariance collapse, physical identifiability collapse, and counterfactual dynamics collapse.

PhyLatent tackles these issues through mechanisms like physical state grounding, future representation alignment, and counterfactual branch separation. Experiments demonstrate that this method significantly reduces the three failure rates and improves Model Predictive Control (MPC) success from 70.0% to 78.1% on OGBench-Cube, while achieving a 98.0% success rate on TwoRooms tasks.

Original post →

More from Embodied

Embodied channel →