Contrastive World Models: dropping pixel reconstruction beats Dreamer in visually complex environments

burny_tech · x · 2026-09-27

A new paper argues that pixel-reconstruction world models waste capacity modeling irrelevant background noise. Building on DeepMind's Dreamer, the authors remove the pixel decoder and instead maximize mutual information between state-action sequences and local patch features of future observations, using a Deep InfoMax-style objective.

Key results:

The authors suggest contrastive, infomax-based objectives are a principled path to world models robust to visual nuisance factors — key for transferring model-based RL agents to the real world.

Related event: Contrastive World Models Drop Pixel Reconstruction and Beat Dreamer(5 posts)→

Original post →

More from Embodied

Embodied channel →