T²Mem lifts robot policy π₀.₅ success rate from 17.93% to 56.83% on memory tasks

yining_hong · x · 2026-10-06

Researchers propose T²Mem, built on a simple insight: memory and action should be learned reciprocally. Rather than a separate module storing history for the policy to consume, memory becomes intrinsic to the policy—action supervision shapes what the agent remembers, and memory shapes how it acts, forming a feedback loop refined around future decisions.

Through test-time training, a single robot policy learns memory representations and memory-conditioned actions from demonstrations alone—no auxiliary memory models, external reasoners, or memory-specific annotations.

Results: on 16 memory-dependent RoboMME tasks, T²Mem raises π₀.₅'s average success rate from 17.93% to 56.83% (+38.90 points). With thanks to Fei-Fei Li and Jiajun Wu.

Original post →

More from Embodied

Embodied channel →