Meta's LSRM wins ECCV 2026 Best Paper Honorable Mention, scales 3D reconstruction with sparse attention
rsasaki0109 · x · 2026-09-30
Meta Reality Labs introduced LSRM (Large Sparse Reconstruction Model), which earned a Best Paper Honorable Mention at ECCV 2026, studying how scaling transformer context windows improves feed-forward 3D reconstruction.
Key finding: expanding the context window — 20× more active object tokens and over 2× more image tokens than prior SOTA — markedly closes the gap with dense-view optimization in fine-grained texture and appearance, enabling high-fidelity 3D object reconstruction and inverse rendering.
Three technical contributions:
- Coarse-to-fine pipeline that predicts sparse high-resolution residuals to focus compute on informative regions;
- 3D-aware spatial routing establishing accurate 2D-3D correspondences via explicit geometric distances instead of attention scores;
- Block-aware sequence parallelism with an All-gather-KV protocol to balance dynamic sparse workloads across GPUs.
On novel-view synthesis benchmarks LSRM delivers >2.4 dB higher PSNR and >40% lower LPIPS than prior SOTA; extended to inverse rendering, its LPIPS matches or beats SOTA dense-view optimization. Paper and code are available.
More from Multimodal
- Flux 3 nails four-way split-screen images of one event from four angles — umesh_ai · 2026-09-30
- MageTrail 2.8B booru finetune costs $593 so far, hits limits of 41k-image dataset — Turbulent-Bass-649 · 2026-09-30
- Filmmakers jam with AI video generation wait times to shoot a music duet with Luma — mrjonfinger · 2026-09-30
- Meta's LSRM Wins ECCV 2026 Honorable Mention, Beats 3D Reconstruction SOTA by 2.4 dB — rsasaki0109 · 2026-09-30
- NVIDIA's LongLive-Plug: Distill Once, Deploy Training-Free Across 54 Downstream Video Models — nvidia · 2026-09-30
- Adobe Research Shows Adversarial Post-Training Restores Missing High-Frequency Detail in Pixel Diffusion — adobe-research · 2026-09-30