LingBot-Map: Open-Source Feed-Forward Model Streams 3D Maps at 20 FPS Over 10,000+ Frames
tom_doerr · x · 2026-09-10
Robbyant open-sourced LingBot-Map (ECCV 2026 oral), a feed-forward foundation model for streaming 3D reconstruction, already at 16.9k GitHub stars.
Key points:
- Geometric Context Transformer: unifies coordinate grounding, dense geometric cues, and long-range drift correction in one streaming framework via anchor context, pose-reference window, and trajectory memory.
- Efficient streaming inference: paged KV cache attention enables stable 20 FPS at 518×378 over sequences exceeding 10,000 frames.
- State-of-the-art reconstruction: outperforms prior methods on benchmarks, building consistent 3D maps from long video streams.
The repo ships the paper PDF, demo scripts, and benchmarks for direct reproduction.
More from Research
- Perplexity releases Q2D-Web: a 190M-document retrieval benchmark for agentic RAG — antoine_chaffin · 2026-09-10
- Correction: the SmolVLM data-filtering trick is from DeepSeek's earlier tech report — eliebakouch · 2026-09-10
- DeepSeek used SmolVLM to quality-filter interleaved pretraining data, researcher spots — eliebakouch · 2026-09-10
- SmolVLM used for strict image-text quality scoring in new model tech report — eliebakouch · 2026-09-10
- DF26 benchmark: humans and detectors near chance on AI speaking videos — Severyn Shykula · 2026-09-10
- Spiral Nd:YAG Waveguides (~100um) Carved via Ion Implantation, FIB and Wet Etch — jwt0625 · 2026-09-10