MetaView: Large-Baseline Synthesis from Single Image

Kwai-Kolors · hf · 2026-07-16

This work introduces **MetaView**, a diffusion-based monocular **novel view synthesis** framework that renders new views under large viewpoint variations from a single input image. The core idea combines two elements: - **Implicit geometric modeling**: Enhances flexibility, avoiding the constraints of overly heavy explicit reconstruction pipelines - **A few but critical explicit 3D cues**: - A feedforward geometry-aware network provides implicit geometric priors to constrain structural consistency - Metric depth anchors the generation to real-world scales The goal is to achieve both **geometric consistency** and **precise controllability**. Experiments show that MetaView outperforms existing methods and generalizes better in the challenging scenario of large-baseline single-image synthesis; the code is open-source.

Related event: Kuaishou Unveils MetaView for Single-Image 3D Synthesis(2 posts)→

Original post →

More from Multimodal

Multimodal channel →