DreamWorld: Geometry-Grounded Two-Stage Video Diffusion for 3D-Consistent World Modeling (ECCV 2026)
机器之心 · wechat · 2026-08-19
HiDream.ai's DreamWorld, accepted to ECCV 2026, tackles unreasonable geometry, object distortion and cross-view inconsistency in camera-controlled video diffusion under large viewpoint changes. It injects spatial priors from a pretrained 3D foundation model and adopts a Geometry-then-Appearance pipeline: a Geometry Video Diffusion completes target-view geometry features from a single image plus camera trajectory via feature-level flow matching, then an Appearance Video Diffusion conditions on those features to synthesize the RGB video.
This geometry-appearance decoupling lets structure reasoning and photorealistic rendering specialize separately. DreamWorld leads Novel View Synthesis on RealEstate10K and Tanks-and-Temples, and scores 75.04 average on WorldScore with stronger 3D consistency.
More from Multimodal
- Opinion: Suno generates listenable sound bites but falls short on full tracks — eschadiol · 2026-08-19
- ComfyUI Workflow: Automated MiniMax Video Segmentation Upscaling — Francky_B · 2026-08-19
- Troubleshooting MiniMax H3 Video-to-Video Character Replacement — NekoBerry420 · 2026-08-19
- Verticals v3 Generates YouTube Shorts for $0.11 Fully Automated — tom_doerr · 2026-08-19
- An AI dream: 100 seamless scenes generated overnight with no human input — wine_dark · 2026-08-19
- JD.com launches Cyber Qixi Gala featuring JoyAvatar digital humans — 京东JoyAI · 2026-08-19