DreamWorld: Geometry-Grounded Video Diffusion for 3D-Consistent World Modeling Accepted at ECCV 2026
jiqizhixin · x · 2026-08-25
DreamWorld is a novel video generation method accepted at ECCV 2026. It splits the process into two stages: generating 3D geometric features using a pretrained 3D foundation model, and then feeding these structural representations into an appearance diffusion model to render final RGB frames. This 'geometry-then-appearance' design provides explicit 3D grounding, avoiding implicit spatial reasoning. It achieves state-of-the-art results on RealEstate10K, Tanks-and-Temples, and WorldScore benchmarks, particularly excelling in large viewpoint changes where other methods suffer from distortions and inconsistencies.
More from Multimodal
- Seeking best cost-effective video avatar model for image+audio input — CelebrationBoth9537 · 2026-08-25
- WAN 3.0 Demo: Multi-shot Choreography and Physics Simulation — minchoi · 2026-08-25
- WAN 3.0 Lands on Pika API with Low-Cost, Realistic Video Generation — minchoi · 2026-08-25
- Redditor makes a full anime trailer with MiniMax H3 I2V/R2V and a 4-step turbo LoRA — Beginning_Tip300 · 2026-08-25
- Flova Introduces Agent-Native Video Workflow Beyond Prompts — Aiden_Tech_Ai · 2026-08-25
- MiniMax H3 multishot test goes wrong with a random Seinfeld guest appearance — Interesting_Room2820 · 2026-08-25