OVIE trains novel-view synthesis on 30M unpaired web images, runs 600x faster than rivals
ducha_aiki · x · 2026-09-09
- The arXiv paper "One View Is Enough!" introduces OVIE, challenging the assumption that novel-view synthesis needs paired multi-view training data.
- Method: a monocular depth estimator serves as a geometric scaffold during training—lifting a source image to 3D, applying a sampled camera transform, and projecting to a pseudo-target view. A masked training formulation restricts geometric, perceptual, and texture losses to valid regions, enabling training on 30 million uncurated internet images.
- At inference OVIE is geometry-free, needs no depth estimator or 3D representation, and outperforms prior methods zero-shot while being 600x faster than the second-best baseline.
- Code and models are publicly available.
More from Multimodal
- Plenty of Cool Image Models, But There's Still Only One Midjourney — umesh_ai · 2026-09-09
- Edinburgh & ESA build COP-GEN, a latent diffusion model that treats Earth observation as a distribution, not a single prediction — anselm · 2026-09-09
- Iterative Music Separation: General Audio Separation Framework Open-Sourced with Training Code — affige_yang · 2026-09-09
- Workflow: photorealistic product UGC with GPT-Image 2.5 via JSON style prompts — sven_ai · 2026-09-09
- Reference-image-to-prompt workflow for photorealistic GPT Images 2.5 outputs — sven_ai · 2026-09-09
- User says they used 'gpt-6 astra' to build a visual essay on the Navier-Stokes problem — paw_lean · 2026-09-09