NAMVIS: next-scale autoregression beats diffusion for multi-view synthesis, 3x faster
Ramil Khafizov · hf · 2026-10-08
- Problem: Sparse-view novel view synthesis is central to 3D content creation, but diffusion approaches are slow due to iterative denoising.
- Method: NAMVIS is diffusion-free, recasting multi-view synthesis as geometry-conditioned next-scale autoregression that predicts discrete visual tokens in a few coarse-to-fine scale steps, sampling tokens in parallel within each scale and across views.
- Geometry anchoring: Multi-scale Projective Pose Encoding injects source/target camera transforms into self- and cross-attention at every resolution; global conditioning plus dense geometry-aware cross-attention preserves source appearance and target-view consistency.
- Results: Outperforms diffusion baselines on PSNR/SSIM/LPIPS across Objaverse, GSO, and OmniObject3D while running over 3x faster.
More from Multimodal
- Musk amplifies user claim that Grok's personalized music is playlist-worthy — elonmusk · 2026-10-08
- Midjourney --sref trick conjures a dramatic '1765 ice cream eating champion' oil painting — dolma33 · 2026-10-08
- Redditor makes a tiny anime episode with WAN 3.0 — tonimestudio · 2026-10-08
- Synthesia's Syren turns one prompt into on-brand video, claiming $20K agency quality for $5 — nikola_mr64990 · 2026-10-08
- obvmp Minimax H3 workflow tested: 10-second clip in ~2 min on an RTX 3090 — fa6637 · 2026-10-08
- Free browser-based tool extracts frame-accurate keyframes from AI videos — QuietPound6205 · 2026-10-08