T²Mem lifts robot policy π₀.₅ success rate from 17.93% to 56.83% on memory tasks
yining_hong · x · 2026-10-06
Researchers propose T²Mem, built on a simple insight: memory and action should be learned reciprocally. Rather than a separate module storing history for the policy to consume, memory becomes intrinsic to the policy—action supervision shapes what the agent remembers, and memory shapes how it acts, forming a feedback loop refined around future decisions.
Through test-time training, a single robot policy learns memory representations and memory-conditioned actions from demonstrations alone—no auxiliary memory models, external reasoners, or memory-specific annotations.
Results: on 16 memory-dependent RoboMME tasks, T²Mem raises π₀.₅'s average success rate from 17.93% to 56.83% (+38.90 points). With thanks to Fei-Fei Li and Jiajun Wu.
More from Embodied
- Real2sim's last mile for robotics: verification pipeline for physics and affordance — chris_j_paxton · 2026-10-06
- Only 30-40k robots installed in the US last year — the case for self-replicating factories — ihorbeaver · 2026-10-06
- ChatGPT Designs Cheap Hexapod Robot and Trains It to Walk in Simulation — lavanyaai · 2026-10-06
- CoRL 2026 Dexterous Manipulation workshop closes submissions tonight — HaozhiQ · 2026-10-06
- Watch, Infer, Coordinate: robots infer a partner's physical limits from watching teamwork, then coordinate zero-shot — mangahomanga · 2026-10-06
- Embodied Analysis launches unified eval platform for physical agents and world models — qinzytech · 2026-10-06