UniWorld-View: Large-Baseline Novel View Synthesis via Video Diffusion
Haiyang Zhou · hf · 2026-08-06
UniWorld-View introduces a unified framework for controllable large-baseline novel view synthesis from monocular inputs.
- The Problem: Reconstruction-based approaches (like NeRF and 3DGS) deteriorate under sparse inputs, while existing generative methods struggle with large-baseline view synthesis due to inaccurate geometric guidance.
- The Method: The framework integrates explicit 3D guidance with generative diffusion modeling. It uses an occlusion-aware point cloud rendering strategy to resolve visibility ambiguities and provide accurate priors for diffusion-based synthesis.
- Performance: By coupling this rendering strategy with powerful video diffusion backbones, it achieves high-fidelity novel view generation even under extreme camera motions and wide-baseline changes. It can also provide multi-view videos for downstream dynamic 3DGS reconstruction.
- Evaluation: Experiments on the WorldScore benchmark and zero-shot NVS benchmarks demonstrate its effectiveness in controllability, geometric consistency, and visual fidelity.
More from Multimodal
- Advanced MiniMax H3 Filmmaking Workflow: From Prompts to Editing — Tricky_Algae2625 · 2026-08-06
- Testing LTX upscaler: Generating high-res images on low VRAM GPUs — AniZeee · 2026-08-06
- Elon Musk Announces Launch of Grok Imagine Image Generation — elonmusk · 2026-08-06
- Meta Unpacks Multimodal Pretraining: Strong Generation with 5% Compute — facebook · 2026-08-06
- ToolArtist: Agentic Image Generation via Unified Multimodal Models — Jiahao Zhao · 2026-08-06
- HelloWorld: Enabling Real-Time Social Interaction in Video World Models — Liangyang Ouyang · 2026-08-06