SPAR3S: sparse autoregressive 3D scene generation from few views, no 3D supervision
kwangmoo_yi · x · 2026-09-05
Researchers from Naver introduce SPAR3S (arXiv:2609.03931), a sparse voxel-aligned 3D latent generative model that completes full 3D scenes from a few unconstrained views, decoded as 3D Gaussian Splats.
- Core idea: formulate scene generation in a compact, structured voxel-aligned latent space where only occupied voxels are represented; completion reduces to predicting missing latent tokens and their spatial support.
- No ground-truth 3D supervision needed: the sparse latent space is learned directly from multi-view images via photometric supervision through differentiable 3D Gaussian Splatting.
- A masked autoregressive transformer jointly models voxel occupancy and latent token values, enabling efficient, spatially consistent generation of unseen regions.
- Unlike feed-forward reconstruction (limited to visible content) or dense volumetric generative models (computationally costly, data-hungry), SPAR3S addresses both limitations at once.
Related event: SPAR3S Generates Full 3D Scenes from Sparse Views Without 3D Supervision(2 posts)→
More from Multimodal
- AI video walks through Sapiens' human history timeline — waterarttrkgl · 2026-09-05
- User locally generates suspense short film with MiniMax H3 ref-to-video in ComfyUI — Fleabum · 2026-09-05
- invideo demos chat-based editing agent that remixes video in 30 minutes — azed_ai · 2026-09-05
- PolyU's OmniColor unifies multi-signal lineart colorization in an ECCV 2026 paper — jiqizhixin · 2026-09-05
- FrankenTTS: pure-Rust Qwen3-TTS port runs voice cloning entirely in your browser — BLUECOW009 · 2026-09-05
- GPT-6 Astra generates Star Destroyers with full interiors, working elevators and hangars — ChrisGPT · 2026-09-05