SplatGuide reuses one 3DGS reconstruction as three priors, hitting SOTA pose-free novel view synthesis

zhenjun_zhao · x · 2026-08-19

New arXiv paper (2608.16863, Yejun Zhang et al.) on photorealistic novel view synthesis from unposed images. Prior pipelines combine feed-forward 3DGS reconstruction with multi-view diffusion but extract at most one signal from the reconstruction — pixel rendering or learned features — and never exploit per-Gaussian visibility for occlusion-aware reference selection, an "information disconnect."

SplatGuide reuses a single 3DGS scene across three complementary roles: rendered images give pixel-aligned geometric conditioning; per-Gaussian source-view indices render into a target-view voting map for occlusion-aware reference selection; reconstruction tokens supply feature-level guidance via cross-attention — all from the same forward pass. It achieves state-of-the-art pose-free novel view synthesis on RealEstate10K, DL3DV, Tanks-and-Temples, and Mip-NeRF 360, and on RealEstate10K with a moderate number of input views it surpasses the ground-truth-pose baseline.

Original post →

More from Multimodal

Multimodal channel →