MemBodied Adds Fixed-Size Episodic Memory to VLA Models, 7.8x Success Rate Gains

Tej Deep Pala · hf · 2026-09-24

Most vision-language-action policies don't retain episode-level information beyond the current observation, hurting history-dependent manipulation; stuffing past observations into context bloats it and adds latency. MemBodied proposes a fixed-size episodic memory with an associative state tracking interactions across policy calls and an episode anchor holding a compact representation of the initial scene, conditioning action generation on inputs plus memory.

Original post →

More from Embodied

Embodied channel →