FactorJEPA: A New World Model for Crowded Urban Environments
Kapil Wanaskar · hf · 2026-08-08
Existing world models are typically evaluated on low-density, lane-structured settings and struggle with populous, chaotic urban environments. To bridge this gap, researchers introduced the DENSEWORLD regime and released a large-scale dataset comprising 1,000 hours of drive-through, walk-through, and aerial video across 22 cities.
The paper also proposes FactorJEPA. Instead of encoding the future in a monolithic latent space, this architecture decomposes scenes into layout, entities, and interactions using separated subspaces and a visibility gate to preserve partially observed agents. Experiments demonstrate improvements in predictive accuracy, causal sensitivity, and robustness, with method rankings replicating consistently across 2B and 1B V-JEPA 2.1 backbones.
More from Research
- NeurIPS 2026 Announces Workshop on User Simulation for AI Evaluation and Training — hyunw_kim · 2026-08-08
- Vision Encoders Rely on Camera Metadata Shortcuts, Explaining AI Image Detection — kwangmoo_yi · 2026-08-08
- The Enduring Value of Data Hinges on the Future Cost of Verification and Generation — oyhsu · 2026-08-08
- HKU's Hengshuang Zhao Named MIT TR35 China for Work in Embodied AI and 3D Vision — YiMaTweets · 2026-08-08
- Genomic Intelligence to Demo DNA Models and Agent Integrations in Upcoming Webinar — julia_kiseleva · 2026-08-08
- AI-Generated Patches Fail Half the Time, Study of 6,000+ Patches Finds — WeldPond · 2026-08-08