MetaView: Large-Baseline Synthesis from Single Image
Kwai-Kolors · hf · 2026-07-16
This work introduces **MetaView**, a diffusion-based monocular **novel view synthesis** framework that renders new views under large viewpoint variations from a single input image. The core idea combines two elements: - **Implicit geometric modeling**: Enhances flexibility, avoiding the constraints of overly heavy explicit reconstruction pipelines - **A few but critical explicit 3D cues**: - A feedforward geometry-aware network provides implicit geometric priors to constrain structural consistency - Metric depth anchors the generation to real-world scales The goal is to achieve both **geometric consistency** and **precise controllability**. Experiments show that MetaView outperforms existing methods and generalizes better in the challenging scenario of large-baseline single-image synthesis; the code is open-source.
Related event: Kuaishou Unveils MetaView for Single-Image 3D Synthesis(2 posts)→
More from Multimodal
- Gemma 4 12B visualization shows what the model predicts from video patches — arjunrajlab · 2026-07-21
- MiniCPM-V 4.6 now runs locally on iPhone with no cloud dependency — amos_gyamfi · 2026-07-21
- Creator says they no longer shoot with a camera, but with prompts — taherdhanera · 2026-07-21
- PixVerse demo turns into a full sci-fi dark comedy set on Mars — aliscodes · 2026-07-21
- Alibaba’s Qwen-Audio-3.0-TTS-Plus takes #1 on Artificial Analysis Speech Arena — airesearch12 · 2026-07-21
- Reddit users say anima_turbo generates brighter images and runs much faster — DeliciousMoraxMora · 2026-07-21