ABot-Recon: Long-Horizon Streaming 3D Reconstruction with 12-Frames Context
zhenjun_zhao · x · 2026-08-31
The paper "Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction" presents ABot-Recon, a streaming model that relies strictly on local context (caching KV features from only the preceding 11 frames). It predicts a point map in the current camera coordinate system and adjacent-frame relative poses. To reduce drift, it uses a lightweight temporal refiner and a composition-aware pose loss. Evaluations show superior long-horizon performance on challenging benchmarks like Oxford Spires.
More from Research
- SenseNova-Vision Formulates Vision as Unified Multimodal Generation — rsasaki0109 · 2026-08-31
- Dietterich: A paper is a structured argument, not a record of how evidence was assembled — tdietterich · 2026-08-31
- LeVJEPA: Preventing video representation collapse with SIGReg — burkov · 2026-08-31
- AI agents spent $3K on research, papers rejected: failure of judgment — rohanpaul_ai · 2026-08-31
- A Lagrangian View of Flow Matching: Why only one step is needed — docmilanfar · 2026-08-31
- NVIDIA's Kimodo Turns Text Into Full-Body 3D Human and Robot Motion — maier_ak · 2026-08-31