RoboMME benchmarks long-horizon memory across 16 robot tasks
micoolcho · x · 2026-07-23
RoboMME targets long-horizon memory in robotics
Long-horizon memory is described as one of the key unsolved problems in robot manipulation, but there have been few good benchmarks to measure progress.
- RoboMME is introduced as a large benchmark for testing how well generalist robot policies can understand language over time.
- It covers 16 robot tasks, including things like object counting and timing-sensitive manipulation.
- The authors evaluate 14 memory-augmented generalist policies across the benchmark suite.
- The goal is to move the field toward a more rigorous understanding of memory as a core robot capability.
The post also points to Episode 91 of RoboPapers with @micoolcho and @chrisjpaxton for more context.
More from Embodied
- Guangdong cookware plant runs lights-out with robots after replacing about 100 workers — heyshrutimishra · 2026-07-23
- Robotaxi sentiment has swung from premature optimism to premature dismissal — skorusARK · 2026-07-23
- Agentic Real2Sim rebuilds a physics simulator from one real-world video — anselm · 2026-07-23
- IIHS says Waymo’s driverless cars crash 68% less often than human drivers — SuzKP · 2026-07-23
- AeroMap3D uses synthetic data to push UAV-to-map localization to 99.2% success — ducha_aiki · 2026-07-23
- Musk says Tesla Semi self-driving should arrive by late this year or early 2026 — XFreeze · 2026-07-23