LoCoSplat ditches learned 3D networks: 4.2x faster, 6.7x less memory feed-forward 3DGS
zhenjun_zhao · x · 2026-10-08
LoCoSplat: feed-forward 3DGS with minimal 3D reasoning
Key insight: a Gaussian is a local primitive—once depth is predicted, the 3D stage only needs local neighborhood statistics, so no heavy learned 3D network is required.
- Point features get a 16-d linear projection into fine+coarse grids, read back per anchor by a 0.14M-parameter pointwise MLP
- No learned 3D network, no dynamic sparse computation; the whole encoder runs as one fp16 CUDA graph
- Outperforms all prior feed-forward methods on PSNR/SSIM/LPIPS at 6/12/24 views on RealEstate10K (+3.3 PSNR over VolSplat at 24 views)
- Reconstructs a 6-view scene in 33ms on one RTX PRO 6000—4.2x faster than prior SOTA, trains 2.7x faster (5.9x at 24 views), uses 6.7x less inference memory
More from Multimodal
- ByteDance's DMAD hits FID 1.04 one-step on ImageNet, SOTA few-step visual generation — ByteDance · 2026-10-08
- Fan-Made Pokemon-Style Brawl MV Created with Vidu Q4 and Suno — RobotMonsterArtist · 2026-10-08
- Luke Wroblewski tests Google's new image model: faster, cheaper, better — LukeW · 2026-10-08
- Envato's Burst mode turns one idea into six image directions for a single AI credit — eptwts · 2026-10-08
- Kandinsky 6.0 launches: 29B Pro and 3B Lite video models with ComfyUI support — KokaOP · 2026-10-08
- 15s MH3 video takes 90 min on a 4080 Super while others claim 8 min on laptops — redpandafire · 2026-10-08