DSReg: Provably Recovering Individual World Latents without Reconstruction
Yujia Zheng, David Klindt, Randall Balestriero, Bernhard Schölkopf
cs.LG, cs.AI, cs.RO, stat.ML
2026-10-07
DSReg turns a rotation-ambiguous JEPA embedding into per-factor latents up to sign, without a decoder. Fifteen pixel seeds match a labeled oracle; dense prediction holds.
A JEPA (joint-embedding predictive architecture: predict the next embedding in representation space, and do not reconstruct pixels) can forecast well while each learned coordinate stays a mixture of several factors. With standard-Gaussian latents and a stationary additive-noise transition, LeJEPA proves that the alignment loss is at least 2(1−ρ)d, where ρ is the correlation between neighboring latent states, with equality if and only if the representation is an orthogonal image of the true state, h(z)=Qz.
An isotropic Gaussian depends on the latent vector only through its length, so a rotation leaves the distribution unchanged. No criterion that sees only that distribution can tell whether cup position and knife position share one coordinate. Dense prediction can stay put while a change to that coordinate drags the knife along. Nonlinear ICA and causal representation learning separate factors with a decoder, a likelihood, auxiliary variables, non-Gaussianity, or interventions. A JEPA has none of those anchors.
The footprint Si of one latent is the set of observed coordinates whose partial derivative is nonzero on a positive-measure set. Structural Diversity asks only that these footprints be pairwise distinct. Overlap is allowed, and so is nesting. A global illumination factor whose footprint contains the local objects fails non-inclusion and structural sparsity, and Structural Diversity still holds. Proposition 1 shows that those older conditions, structural variability included, are strictly stronger. Two latents with the same footprint cannot be separated by the support pattern alone.
LeJEPA fixes the orthogonal class. DSReg (dependency-sparsity regularization) then searches for one more rotation that makes the dependency support as sparse as possible. Theorem 2 adds functional no-cancellation: active partials must not cancel each other. Every minimizer is a signed permutation, so each estimated coordinate equals one true latent times ±1.
The encoder stays frozen.
The ℓ1 penalty scores magnitude as well as the number of nonzeros, so it can prefer a different rotation from an exact support count. On the Figure 8 benchmarks the exact count does at least as well, and refining the ℓ1 solution with that count closes the gap. The regressions train no decoder and do not optimize reconstruction. A post-hoc fit and joint training recover equally well; the joint head remains slightly behind until the decoupled fit finishes (Appendix Table 2).
On one 48 GB GPU the procedure reaches latent dimension 8192, 10^6 samples, and observation dimension 2^17, with per-step cost essentially flat. Theorem 3 gives the noise margin: if the threshold plus Jacobian error stays below the product of signal strength and pattern margin, minimizers land near a signed permutation.
Dense R² checks whether the full state is still linearly spanned. MCC checks one-to-one recovery. The baseline is always the same representation before the rotation. Visual encoders use 5 to 15 seeds; the other settings use 10 to 20.
| Setting | Baseline | DSReg |
| Nested synthetic footprints, 20 runs | LeJEPA stays mixed | Near ceiling at every dimension |
| Identical footprints | Individual recovery impossible | Pair-span CCA of 1.00 |
| Pixel encoders, 15 seeds | PCA, Varimax, FastICA | Matches the best labeled orthogonal alignment |
| True dim 8, estimate widened to 32 | Same labeled alignment | Gap within 0.001, 5 seeds |
The synthetic footprints are nested, which is exactly the regime earlier structural conditions rule out. On that same estimated representation, PCA, Varimax, and FastICA all fail to unmix the factors (Appendix Table 5). The signal that works is the Jacobian dependency. Figure 3b checks the regimes in Proposition 1 on analytic orbits h=Qz at N=16. Wherever footprints differ, recovery is near perfect, and the nested regime is as clean as the easy one. Identical footprints retain only the shared subspace.
Six probes cover visual editing, sparse control, short-horizon prediction, and surprise detection, across TwoRoom, FetchSlide, and PushT. DSReg improves every probe and wins on nearly every run, while the dense span stays matched (Appendix Table 6). After encoders are trained from pixels, it leads on every run of all three probes. Gaussian 3DShapes rises from a mixed representation to near-ceiling recovery at a matched dense R². On quarter-orientation dSprites the linear-identifiability premise is only partly met; DSReg still beats swept β-VAE and β-TCVAE baselines, and that premise is the ceiling. Per-probe margins are not printed in the text.
A trained checkpoint can stay frozen. The new piece is one orthogonal head, and dense readouts do not have to be retrained. Control, editing, and monitoring that should move one factor at a time switch onto the rotated coordinates.
Nested footprints remain identifiable under this condition. The paper calls the result the first JEPA that recovers every world latent without reconstruction, labels, or a non-Gaussian assumption.
The orthogonal class comes from LeJEPA on a Gaussian world. When that premise is incomplete, no clean class is left to search.
The benchmarks are simulated or rendered. There is no experiment on a physical robot. Factors that share a footprint cannot be separated by support alone; the discussion places the fix in temporal structure or cheap interventions. The Gaussian world is a foothold, and the guarantee extends only as far as linear identifiability does.
The default objective is the ℓ1 relaxation. It can select a different rotation from the exact support count, and the evidence that the count does at least as well is Figure 8, not the default setup. Theorem 3 covers the thresholded population criterion. The main text has no noise table for the finite-sample estimator. Whether two factors in natural video really leave different footprints is unchecked. The observation vector used for the Jacobian is chosen before fitting, and the encoder is allowed to drop low-level pixel detail.