GAE: A Geometry-Native Autoencoder Cuts World Model FVD by 23.1%
yshan2u · x · 2026-09-22
GAE (Geometry-native Autoencoder) is introduced as a foundational building block for persistent world models: a compact geometry-native autoencoder is built directly into the world modeling process, giving generated worlds a 'soft skeleton' that stays coherent as the camera moves. Trained from scratch, the resulting world model reduces FVD by up to 23.1% and halves camera-trajectory error.
More from Research
- SVEET: streaming video editing with a diffusion model hits 15 FPS on a single H100 — SJTU · 2026-09-22
- Video Summagator tool revives a 2012 CHI paper's space-time cube technique — mishig25 · 2026-09-22
- Semantic cache feedback loop quietly breaks its own guaranteed 2% error rate, worst case 3x worse — Reasonable_Royal_621 · 2026-09-22
- Stanford turns papers into AI agents; two unrelated studies surface unreported ADHD genetic link — VraserX · 2026-09-22
- Xiaomi MiMo 2.6 Pro formalizes Li–Yorke chaos theorem in 6,000+ lines of verified Lean — bookwormengr · 2026-09-22
- Apollo, the first advanced LLM for Ancient Greek, aims to restore tattered papyri — nordicinst · 2026-09-22