Contrastive World Models: dropping pixel reconstruction beats Dreamer in visually complex environments
bonniesjli · x · 2026-09-26
Bonnie Li released an arXiv paper on Contrastive World Models, a way to learn latent dynamics without pixel reconstruction:
- Method: Built on Dreamer, it replaces observation reconstruction with a Deep InfoMax-style lower bound that maximizes mutual information between state-action sequences and local patch features of future observations, so representations keep only planning-relevant information.
- Results: In small-scale experiments, it matches Dreamer and a momentum-prediction baseline in default settings, and substantially outperforms both once distractors or natural video backgrounds are introduced, while training faster since the pixel decoder is removed entirely.
- Takeaway: Written independently between roles with minimal compute; infomax-based objectives look like a principled path to world models robust to visual nuisance factors—key for transferring model-based RL agents to the real world.
Related event: Contrastive World Models Train World Models Without Pixel Reconstruction(3 posts)→
More from Embodied
- Musk lays out his 90%-likely AI future: personal robots and universal high income — XFreeze · 2026-09-26
- Why robots are humanoid: the models are trained on humans doing the work — floguo · 2026-09-26
- Tesla workers balk at training Optimus as Fremont plant stops Model S/X production — Ars Technica AI · 2026-09-26
- Trader on AI capex: finance people slow to catch Kimi/GLM coding-agent signals; robotics still pre-PMF — menhguin · 2026-09-26
- Musk: Human bandwidth is ~100 bits/sec, and Neuralink aims to close the gap with terabit AI — XFreeze · 2026-09-26
- AprilTag spotted at a Walmart in NY turns out to guide BrainCorp's autonomous floor scrubbers — carlosdponx · 2026-09-26