GAE: A Geometry-Native Autoencoder Cuts World Model FVD by 23.1%

yshan2u · x · 2026-09-22

GAE (Geometry-native Autoencoder) is introduced as a foundational building block for persistent world models: a compact geometry-native autoencoder is built directly into the world modeling process, giving generated worlds a 'soft skeleton' that stays coherent as the camera moves. Trained from scratch, the resulting world model reduces FVD by up to 23.1% and halves camera-trajectory error.

Original post →

More from Research

Research channel →