Latent-Foresight: end-to-end latent world models beat two-stage pipelines on future scene prediction
Efstathios Karypidis · hf · 2026-10-02
Latent-Foresight is an end-to-end framework that jointly learns a latent tokenizer and a flow-based generative dynamics model, explicitly shaping representations for temporal predictability in world modeling. Unlike two-stage pipelines that compress VFM features with fixed dimensionality reduction before training a separate predictor, it guarantees the latent space supports predictable dynamics, with design choices preventing latent collapse and aligning reconstruction with generative objectives. It learns more temporally coherent representations, consistently outperforms two-stage baselines across future scene understanding tasks and horizons, and removes separate training stages even during high-resolution adaptation. Code and weights are open-sourced on GitHub.
More from Research
- LOCI: hybrid spatial linear memory lets streaming world models recall revisited scenes at ~30% less memory — IFM · 2026-10-02
- SAKIKO auditing shows +55 net-gain interventions corrupt over half of correct tool-using LLM decisions — UniversityofBirmingham · 2026-10-02
- Researchers warn AI-written 'salami' papers are flooding arXiv with low-value work — EhudReiter · 2026-10-02
- Arena Physica explains why FEM solvers never compute E-fields at mesh nodes — burny_tech · 2026-10-02
- AWSM grounds LLM-agent 3D scene reconstruction in IMU, depth and pose evidence — anselm · 2026-10-02
- Study links attention nonlinearity to power-law massive activations and scaling laws — burny_tech · 2026-10-02