DreamWorld: Geometry-Grounded Video Diffusion for 3D-Consistent World Modeling Accepted at ECCV 2026

jiqizhixin · x · 2026-08-25

DreamWorld is a novel video generation method accepted at ECCV 2026. It splits the process into two stages: generating 3D geometric features using a pretrained 3D foundation model, and then feeding these structural representations into an appearance diffusion model to render final RGB frames. This 'geometry-then-appearance' design provides explicit 3D grounding, avoiding implicit spatial reasoning. It achieves state-of-the-art results on RealEstate10K, Tanks-and-Temples, and WorldScore benchmarks, particularly excelling in large viewpoint changes where other methods suffer from distortions and inconsistencies.

Original post →

More from Multimodal

Multimodal channel →