LEDGER builds persistent 3D object memory from egocentric video, answers questions without rewatching

mangahomanga · x · 2026-10-08

Researchers at Johns Hopkins introduce LEDGER, a persistent 3D object memory built from egocentric video: it detects objects in sampled frames, localizes them in 3D via camera pose, and links repeated sightings over time — objects persist in memory even after leaving view.

Each object's history is split into rest segments; a move is only recorded when multiple observations agree, and segments carry short descriptions (e.g. container contents). Questions are answered later from these text records without rewatching footage. In a demo, three separate kitchen recordings were stitched into one 21-minute stream and LEDGER built a unified memory. Across 100 streams, revisiting a scene cost little accuracy, but scene changes lowered it; per-scene memories recovered part of the loss. Paper coming to arXiv, code to be released.

Related event: LEDGER Builds Persistent 3D Object Memory from Egocentric Video(2 posts)→

Original post →

More from Embodied

Embodied channel →