SplatGuide reuses one 3DGS reconstruction as three priors, hitting SOTA pose-free novel view synthesis
zhenjun_zhao · x · 2026-08-19
New arXiv paper (2608.16863, Yejun Zhang et al.) on photorealistic novel view synthesis from unposed images. Prior pipelines combine feed-forward 3DGS reconstruction with multi-view diffusion but extract at most one signal from the reconstruction — pixel rendering or learned features — and never exploit per-Gaussian visibility for occlusion-aware reference selection, an "information disconnect."
SplatGuide reuses a single 3DGS scene across three complementary roles: rendered images give pixel-aligned geometric conditioning; per-Gaussian source-view indices render into a target-view voting map for occlusion-aware reference selection; reconstruction tokens supply feature-level guidance via cross-attention — all from the same forward pass. It achieves state-of-the-art pose-free novel view synthesis on RealEstate10K, DL3DV, Tanks-and-Temples, and Mip-NeRF 360, and on RealEstate10K with a moderate number of input views it surpasses the ground-truth-pose baseline.
More from Multimodal
- AI Tool GeoSpy Locates Photos from Pixels with Meter-Level Accuracy — saibharadwaj · 2026-08-20
- Runway Gen-2 Update: 1080p Support, 50 References, 30s Generation via API — tlakomy · 2026-08-20
- Digital Sculpting: Creating the Thesis Rock with Rendering Magic — every · 2026-08-20
- Swarms Builds Inference Engines for Media Generation without Frameworks — bingxu_ · 2026-08-20
- AI places famous internet memes on a single street — _jaydeepkarale · 2026-08-20
- User shares fun generated results using Kling AI Omni 3 — LudovicCreator · 2026-08-20