LEDGER builds persistent 3D object memory from egocentric video, answers questions without rewatching
mangahomanga · x · 2026-10-08
Researchers at Johns Hopkins introduce LEDGER, a persistent 3D object memory built from egocentric video: it detects objects in sampled frames, localizes them in 3D via camera pose, and links repeated sightings over time — objects persist in memory even after leaving view.
Each object's history is split into rest segments; a move is only recorded when multiple observations agree, and segments carry short descriptions (e.g. container contents). Questions are answered later from these text records without rewatching footage. In a demo, three separate kitchen recordings were stitched into one 21-minute stream and LEDGER built a unified memory. Across 100 streams, revisiting a scene cost little accuracy, but scene changes lowered it; per-scene memories recovered part of the loss. Paper coming to arXiv, code to be released.
Related event: LEDGER Builds Persistent 3D Object Memory from Egocentric Video(2 posts)→
More from Embodied
- Comma_ai user praises driving data setup, but car chewed through the cables — fforres · 2026-10-08
- ~200 production Cybercabs without steering wheels spotted staging in Houston — JOBhakdi · 2026-10-08
- NavSafe-∞: photorealistic closed-loop benchmark exposes safety gap across 20 E2E driving policies — zhoubolei · 2026-10-08
- SEAR benchmark tests how LLM agents self-evolve on robots, from DeepSeek to Astra — xiye_nlp · 2026-10-08
- Polymarket bets open on whether Meta ships its Tamagotchi-style Muse Charm AI device by December — Polymarket · 2026-10-08
- Microsoft's $5,999 Surface RTX Spark Dev Box preorders open, ships November — tomwarren · 2026-10-08