SPAR3S generates complete 3D scenes from sparse views via sparse voxel-aligned autoregressive modeling

zhenjun_zhao · x · 2026-09-05

A Naver Labs Europe paper introduces SPAR3S, a sparse voxel-aligned 3D latent generative model for conditional scene completion from unconstrained multi-view images, requiring no ground-truth 3D supervision. The latent space — representing only occupied voxels — is learned from images via photometric supervision through differentiable 3D Gaussian Splatting. A masked autoregressive transformer jointly predicts voxel occupancy and latent tokens, enabling spatially consistent generation of unseen regions, validated on synthetic indoor scenes.

Related event: SPAR3S Generates Complete 3D Scenes from Sparse Views(3 posts)→

Original post →

More from Multimodal

Multimodal channel →