ModAR: First Multimodal Autoregressive World-Action Model Cuts Training Compute 20x
Adam Hung · hf · 2026-09-16
ModAR is the first world-action model (WAM) to autoregressively denoise multiple future modalities before predicting actions, letting each prediction condition on previously generated modalities.
- Key finding: predicting point tracks, DINO features, and depth maps helps WAMs; additionally predicting future RGB offers no consistent benefit.
- Performance: ModAR's sequential generation outperforms existing WAM formulations with the highest average success rate at all data scales.
- Efficiency: fine-tuned on the video-initialized Flex-π, ModAR achieves a slightly higher success rate (75% vs 72%) with 20x fewer training FLOPs and no pretraining.
- Real world: outperforms baselines on three real bimanual tasks and improves with human videos.
More from Embodied
- Driverless trucks hit Ohio roads: L3 on interstates, L4 piloting at 15 mph — bennash · 2026-09-16
- Semi-humanoid robot OpenShell debuts at exhibition, dubbed "sloth robot" — chris_j_paxton · 2026-09-16
- Flow Chaser brings on-device AI level generation to Apple Vision Pro rhythm gaming — Scobleizer · 2026-09-16
- Watch: how a farming robot sees the field — Scobleizer · 2026-09-16
- JPMorgan sees 25M+ GPU/ASIC shipments by 2028, ASICs dominate — a 'narrative violation' — bookwormengr · 2026-09-16
- OpenAI and Apple both reportedly exploring Pixar-lamp-style robot prototypes — flowersslop · 2026-09-16