Ant Group's LingBot-Map, ECCV 2026 Oral: streaming 3D reconstruction from one RGB camera at ~20 FPS
jiqizhixin · x · 2026-09-20
Ant Group's Lingbo (Robbyant) team presents LingBot-Map, selected as ECCV 2026 Oral. The paper "Geometric Context Transformer for Streaming 3D Reconstruction" builds a spatial map incrementally the way humans do when entering a new room: from a single ordinary RGB camera it estimates camera pose and reconstructs 3D scene structure in real time during video capture. It runs streaming reconstruction at 20 FPS on a single GPU, handles sequences beyond 10,000 frames, and stays robust across multi-room traversals, dramatic environment changes and large viewpoint shifts — a new route to "how robots remember space" without complex hardware or offline processing.
More from Embodied
- World Labs CEO Alex Kendall to Headline Long Horizon Physical AI Summit — alexgkendall · 2026-09-20
- Dev shows robot arm controlled by Claude, quips robotics is 'far too easy' and returns to GPU kernels — KuterDinel · 2026-09-20
- Sitzmann clarifies: only learned systems generalize to unknown in-the-wild objects — vincesitzmann · 2026-09-20
- Neuralink releases 'Speaking With The Mind', a 7-minute film on its BCI — Competitive_Travel16 · 2026-09-20
- AetherAI's CausalWM tops TriWorldBench with causal chain-of-thought world modeling — 机器之心 · 2026-09-20
- GPT-6 Astra drives a $150 robot arm to paint, hits 95% on block pickup — 机器之心 · 2026-09-20