Ant Group's LingBot-Map, ECCV 2026 Oral: streaming 3D reconstruction from one RGB camera at ~20 FPS

jiqizhixin · x · 2026-09-20

Ant Group's Lingbo (Robbyant) team presents LingBot-Map, selected as ECCV 2026 Oral. The paper "Geometric Context Transformer for Streaming 3D Reconstruction" builds a spatial map incrementally the way humans do when entering a new room: from a single ordinary RGB camera it estimates camera pose and reconstructs 3D scene structure in real time during video capture. It runs streaming reconstruction at 20 FPS on a single GPU, handles sequences beyond 10,000 frames, and stays robust across multi-room traversals, dramatic environment changes and large viewpoint shifts — a new route to "how robots remember space" without complex hardware or offline processing.

Original post →

More from Embodied

Embodied channel →