MM-ABC robot foundation model hits 83% on real-world mobile manipulation tasks
Qiwei Liang · hf · 2026-10-07
MM-ABC is a foundation model for mobile manipulation built around Seeing, Coordinating, and Imagining arm-base collaboration.
- Uses sparse multi-level VLM features for spatial perception, a training-only future branch leveraging world imagination and geometric intent as extra supervision, and MM-APT to coordinate separate manipulation and mobility streams via masked joint attention and clean-action x-prediction
- Ablations: replacing clean-action prediction with velocity prediction drops RoboCasa365 composite-seen success from 32.8% to 29.2%; removing future supervision or multi-level conditioning hurts more
- Pretrained on 5,000+ hours of heterogeneous robot data spanning 400K+ episodes, 12 datasets, and 17 embodiments
- Results: 44.71% on EBench, 61.2% on RoboCasa365, 99.1% on LIBERO, 82.8% on LIBERO-Plus, and 83% mean success on five real-world tasks
More from Embodied
- Tesla's Cybercab rests on aggressive FMVSS interpretation NHTSA could reject — binarybits · 2026-10-07
- Waymo Colors Inside FMVSS Lines; Tesla Bets on Regulatory Favor, Analyst Argues — binarybits · 2026-10-07
- 19-Year-Old Founder Zain Raises $11M Seed Led by a16z to Ship Personal AI Computer Core — nick_linck · 2026-10-07
- Father-son team launches Rhem, a companion robot helping aging parents call, book and track health — ycombinator · 2026-10-07
- Sentdex finds decision language models fall short in real robotics pipelines vs RL/VLAs — Sentdex · 2026-10-07
- NeurIPS paper adds probabilistic uncertainty quantification to robot memory for better retrieval — lucacarlone1 · 2026-10-07