DreamWorld: Geometry-Grounded Two-Stage Video Diffusion for 3D-Consistent World Modeling (ECCV 2026)

机器之心 · wechat · 2026-08-19

HiDream.ai's DreamWorld, accepted to ECCV 2026, tackles unreasonable geometry, object distortion and cross-view inconsistency in camera-controlled video diffusion under large viewpoint changes. It injects spatial priors from a pretrained 3D foundation model and adopts a Geometry-then-Appearance pipeline: a Geometry Video Diffusion completes target-view geometry features from a single image plus camera trajectory via feature-level flow matching, then an Appearance Video Diffusion conditions on those features to synthesize the RGB video.

This geometry-appearance decoupling lets structure reasoning and photorealistic rendering specialize separately. DreamWorld leads Novel View Synthesis on RealEstate10K and Tanks-and-Temples, and scores 75.04 average on WorldScore with stronger 3D consistency.

Original post →

More from Multimodal

Multimodal channel →