MemBodied Adds Fixed-Size Episodic Memory to VLA Models, 7.8x Success Rate Gains
Tej Deep Pala · hf · 2026-09-24
Most vision-language-action policies don't retain episode-level information beyond the current observation, hurting history-dependent manipulation; stuffing past observations into context bloats it and adds latency. MemBodied proposes a fixed-size episodic memory with an associative state tracking interactions across policy calls and an episode anchor holding a compact representation of the initial scene, conditioning action generation on inputs plus memory.
- On five memory-demanding RMBench tasks: 7.81x mean success rate over stateless policy, 2.98x over vanilla recurrent memory, 1.3x over the strongest memory baseline with 10x fewer added parameters
- 90.6% on fully observable LIBERO-Long, a 5.4% improvement over stateless π0
- A practical alternative to growing policy context for history-dependent manipulation
More from Embodied
- Unverified: Meta rumored to unveil Muse Charm voice keychain agent device — davidpattersonx · 2026-09-24
- Danfei Xu: "Robots are just computers" — a deflating realization — fdellaert · 2026-09-24
- Tesla Robotaxi in Austin rides smoother than any Uber or Lyft, user reports — EricETesla · 2026-09-24
- GPT-6 with a robot arm executed most of 5 dangerous tasks: stabbing a dummy, tossing gas canisters into a furnace — FuSheng_0306 · 2026-09-24
- Huawei Watch D3 unboxed: £399.99 smartwatch with inflatable cuff for medical-grade blood pressure — CurieuxExplorer · 2026-09-24
- Shanghai AI Lab's InternW0 Is a Physical World Model Unifying Video Prediction and Robot Control — Jisong Cai · 2026-09-24