Preserving Unstable Modes Through Inverse Dynamics in JEPA World Models
Leonardo F. Toso, Yann LeCun, James Anderson, Oumayma Bounou
cs.LG, cs.RO, eess.SY, math.OC
2026-10-06
Next-step JEPA plus SIGReg can minimize its loss while dropping every unstable mode; EP-IDM lifts CartPole latent LQR success from 0% to 100%.
Legged robots and quadrotors sit near unstable equilibria. Leave a small tilt uncorrected and it grows without bound. Visual control usually compresses pixels into a latent state and plans there. A joint-embedding predictive architecture (JEPA) skips pixel reconstruction: an encoder emits a latent, a predictor steps it forward, and an anti-collapse regularizer stops every observation from landing on one constant. The regularizer used here is SIGReg, which pulls the batch of latents toward an isotropic Gaussian.
A spread-out latent cloud does not mean the directions feedback needs are still there. An unstable mode is a state direction whose eigenvalue has magnitude at least 1. Balance a pole: if the tilt is absent from the representation, feedback cannot catch it. Lemma 1 is the sharp condition. A causal controller that sees only latents can asymptotically stabilize the linear system if and only if the encoder keeps every unstable mode and every marginally stable mode. That property is detectability.
Lemma 2 shows the standard loss can miss it. Split the state into an unstable block and a stable block. An encoder that reads only the stable block, with a predictor that copies the stable subsystem, drives one-step error to zero. Whiten the stable coordinates and the latent is standard normal, so SIGReg is minimized as well. For every regularizer weight the training objective can sit at its minimum while the whole unstable subspace lies in the encoder's kernel.
The counterexample is a linear encoder with little compression, so capacity is not the excuse. On a synthetic linear system and on linearized CartPole (Figure 2), the dominant direction learned with SIGReg misses the true unstable eigenvector and latent LQR diverges. EP-IDM aligns that direction, and the same LQR brings trajectories back.
The added term is endpoint inverse dynamics (EP-IDM). Prediction remains, either one-step (1SP) or autoregressive multi-step (MSP), and matches predicted latents to encoded future latents. EP-IDM sees only the first and last latent in a window and reconstructs the action sequence between them. It replaces SIGReg. Two other variants reconstruct each action from adjacent latents, or the full sequence from the whole latent trajectory.
In a linear system, the directions an action sequence can reach within H steps span the reachable subspace. Theorem 1: if, for almost every initial state, the conditional support of the action sequence contains an open set, and if an affine inverse model drives its loss to zero, the encoder is injective on that subspace. No direction an H-step action sequence can produce is discarded. When the plant is stabilizable, the raw observations are observable, and H is long enough, the unstable subspace sits inside the reachable subspace, so detectability holds. Exact recovery also needs the stacked action dimension to be no larger than the latent dimension. Past that point the loss is only a soft bias, which is how the experiments use it.
The finite-horizon controllability Gramian sums the energy that inputs inject into each direction. Making H longer stops adding new reachable directions, but unstable responses grow exponentially while stable ones shrink, so the leading eigenspace converges to the unstable subspace (Theorem 2). The argument requires controllability energy that grows at least as α^{2H} with α > 1, and coupling through the stable block must not cancel that response. Eigenvalues on the unit circle are excluded.
The encoder takes two consecutive frames, their difference, and proprioception. On CartPole, LQR linearizes the predictor at the equilibrium. CEM searches sampled rollouts; GBP backpropagates a terminal latent cost. Both replan on a receding horizon. Walker2D uses iCEM.
Latent LQR is the direct test. If the local linearization dropped the unstable mode, this feedback cannot hold the pole. Success on CartPole means the state norm is at most 0.7 after 300 steps. MFS is the fraction of steps inside that band.
| Method | LQR success | MFS | CEM | GBP |
| Ground truth | 100% | 1.000 | 100% | 100% |
| 1SP + SIG | 0% | 0.202 | 100% | 90% |
| MSP + SIG | 0% | 0.247 | 100% | 100% |
| 1SP + EP-IDM | 100% | 0.999 | 100% | 100% |
| MSP + EP-IDM + SIG | 100% | 0.996 | 100% | 100% |
In Figure 4 the SIG state norm grows without bound. EP-IDM returns it to the equilibrium and holds it there. CEM and GBP still score 90% to 100% with SIG models. They search finite nonlinear rollouts and replan, so they do not need the unstable mode inside the local linearization.
Walker2D asks whether a walking limit cycle survives. A SAC policy rolled for 500 closed-loop steps keeps a periodic right-hip phase portrait under 1SP+EP-IDM. Both SIG runs smear it, and only EP-IDM still looks like walking once latents are decoded to pixels. Under latent iCEM, averaged over ten trials, EP-IDM reaches 2.91 m/s and 9.15 m. Ground truth is 3.61 m/s and 14.43 m. 1SP+SIG manages 0.08 m/s and 0.06 m; MSP+SIG manages 0.46 m/s and 0.66 m. Final torso height is 0.92 m for EP-IDM and 1.21 m for the true dynamics.
PointMaze is an open-loop stable point mass in a U-maze, a check that inverse dynamics does not break tasks that need no stabilization. Action-reconstruction variants trained from scratch reach at least 90% CEM and at least 70% LQR. MSP+SIG scores 90% CEM, 60% LQR, and 100% GBP. Ground truth is 100% CEM, 80% LQR, and 90% GBP. GBP estimates gradients with SPSA, two rollouts per step, because MuJoCo is not differentiable. Frozen DINOv2 with MSP falls to 50% CEM and 20% LQR. iBOT with MSP scores 100% CEM, 80% LQR, and 80% GBP.
Figure 9: LQR from 1SP+EP-IDM stabilizes almost every tested initial state, and the Lyapunov difference is negative near the upright equilibrium. MSP+SIG covers only a small neighborhood.
A low prediction loss and a non-constant latent are the usual JEPA signal that a representation is ready for control. On an open-loop unstable plant that signal is false. The loss can be minimal, and the latents need not collapse on the training distribution, while local feedback still cannot see the mode it has to correct.
EP-IDM is one extra head: reconstruct the action sequence from the endpoint latents. On CartPole it moves latent LQR from 0% to 100%, and MFS from 0.202 and 0.247 to 0.996 and 0.999, against 1.000 for the true dynamics. For visual stabilization that depends on linear feedback near an equilibrium, anti-collapse regularization does not replace this constraint.
On PointMaze, SIG already has 90% CEM. On CartPole, SIG also succeeds at CEM and GBP. The failure sits in local linear feedback. Sampling-based replanning can route around it. What is added is a detectability condition on JEPA training, not a new planner.
The theorems stop at linear systems and affine encoders. The image tasks show the same pattern on nonlinear CartPole, Walker2D, and PointMaze, without a nonlinear proof. Zero inverse-dynamics loss needs rich action excitation and a stacked action dimension no larger than the latent size. When the latent is smaller, the experiments use a soft penalty, which leaves a gap from Theorem 1. The α^{2H} energy lower bound, the requirement that coupling not cancel the unstable response, and the exclusion of unit-circle eigenvalues are not checked on the visual tasks.
Walker2D displacement is 9.15 m against 14.43 m for the true dynamics, and final height is 0.92 m against 1.21 m. The phase portraits go through an MLP that decodes latents back to physical state, a different path from pure latent control. The terminal-distance cost is described as deliberately simple, and Walker planning adds a decoder.
Reported success rates use ten trials. SIG models already succeed at CEM and GBP on CartPole, so a pipeline that only replans by sampling may never see the missing mode. Frozen DINOv2 on PointMaze lags both from-scratch training and iBOT, and that gap is left unexplained.