New theory shows predictive self-supervised learning provably separates stochastic signals from distractors
Friedemann Zenke's team, led by Fabian Mikulasch, released an arXiv paper, "Predictive Self-Supervised Learning Provably Identifies Stochastic Signals…", accompanied by a multi-post researcher thread on Twitter. The paper theoretically proves that latent-space predictive self-supervised methods (JEPA, CPC, SimCLR-style) can identifiably recover stochastic signals under distracting noise, validated with MuJoCo inverted pendulum experiments.
Confirmed
- The thread frames the core paradox: the signals we care about are themselves stochastic, and both signal and nuisance variables make the next observation unpredictable — so how does a predictive model know what to keep and what to ignore? The authors note that existing identifiability theories cover either stochastic signal dynamics (without nuisance variables) or nuisance variables only (requiring deterministic signal dynamics), whereas video, sensors, and agents interacting with the world must handle both at once — precisely the gap this paper fills.
- Theoretical result: common SSL methods separate signal from nuisance via two cooperating mechanisms — predictive mutual information maximization retains all predictable information, while latent distribution matching makes the retained signal identifiable; with Gaussian predictors, recovery up to affine equivalence is provable.
- Experimental setup and results: on a MuJoCo hopper with non-deterministic dynamics plus randomized colors, lighting, cameras, and noisy backgrounds, representations from latent predictive models support linear decoding of poses and velocities; however, adding a pixel reconstruction loss to the same model causes the learned latents to partially encode nuisance information, no longer permitting linear decoding of the system state.
- Ongoing work shows the learned latent dynamics can be rolled out forward from a few initial observations (t<0), recovering the hopper's signals along with uncertainty estimates.
Why it matters
- This work elevates the common empirical observation that "JEPA-style methods are robust to changes in lighting, viewpoint, and background" into a provable identifiability theory, explicitly covering the setting of "stochastic signals coexisting with nuisance variables" that prior theories couldn't handle — directly relevant to video modeling and embodied agent learning.
- The contrast showing pixel reconstruction loss harms linear decodability provides both theoretical and experimental support for choosing predictive objectives over reconstruction objectives.
2026-10-08 ~ 2026-10-08 · 8 related posts
Primary sources
- New paper proves predictive SSL provably identifies stochastic signals under nuisance — hisspikeness ·
- New theory shows SSL separates stochastic signals from nuisance via MI maximization plus distribution matching — hisspikeness ·
- MuJoCo hopper experiment validates latent predictive representations decode pose and velocity linearly — hisspikeness ·
- New thread unravels why JEPA-style latent prediction works: an identifiability theory for SSL — hisspikeness · 2026-10-08
- The conundrum of latent prediction: signals themselves are stochastic, so what to keep? — hisspikeness · 2026-10-08
- Prior identifiability theory can't handle stochastic signals and nuisance together, argues new thread — hisspikeness · 2026-10-08
- [source] New theory shows SSL separates stochastic signals from nuisance via MI maximization plus distribution matching — hisspikeness · 2026-10-08
- [source] MuJoCo hopper experiment validates latent predictive representations decode pose and velocity linearly — hisspikeness · 2026-10-08
- Adding a pixel-reconstruction loss makes latents unable to decode hopper state — hisspikeness · 2026-10-08
- Learned latent dynamics roll forward from few observations to recover hopper state with uncertainty — hisspikeness · 2026-10-08
- [source] New paper proves predictive SSL provably identifies stochastic signals under nuisance — hisspikeness · 2026-10-08