Contrastive World Models: latent-space world models beat pixel reconstruction under distractors

bonniesjli · x · 2026-09-26

bonniesjli presents Contrastive World Models: train latent states to maximize mutual information with future observations via Deep InfoMax, dropping pixel reconstruction entirely. In small-scale experiments CWM matches pixel-reconstruction and momentum-prediction baselines by default, and substantially outperforms both when distractors or natural video backgrounds are added, while training more efficiently without a decoder.

Related event: Contrastive World Models Train World Models Without Pixel Reconstruction(3 posts)→

Original post →

More from Research

Research channel →