ModAR Challenges RGB Representations with 30M Parameters, Beating 6B Models
Researchers introduce ModAR, the first world-action model that autoregressively denoises multiple future modalities before predicting actions. With only 30M parameters trained from scratch, it outperforms fine-tuned 6B-parameter video models while using 20x less compute.
2026-09-16 ~ 2026-09-17 · 2 related posts
- ModAR: First Multimodal Autoregressive World-Action Model Cuts Training Compute 20x — Adam Hung · 2026-09-16
- ModAR: a 30.1M robot world model trained from scratch beats a 6B video-model baseline — CSProfKGD · 2026-09-17