ModAR: a 30.1M robot world model trained from scratch beats a 6B video-model baseline
CSProfKGD · x · 2026-09-17
Adamjhung's team questions whether RGB pixels are the right representation for robot world-action models, arguing RGB wastes capacity on fine-grained details irrelevant to robot policies.
- DINO features, point tracks, and depth capture semantics, motion, and geometry — but no single modality suffices.
- ModAR predicts the future one modality at a time, with each prediction informing the next, trained from scratch.
- The 30.1M scratch-trained model outperforms a 6B video-model-initialized baseline fine-tuned for robotics.
Related event: ModAR Challenges RGB Representations with 30M Parameters, Beating 6B Models(2 posts)→
More from Embodied
- Two ways to run GPT-6 Astra on robots: direct VLA policy or agent with tools — YuXiang_IRVL · 2026-09-17
- Embodied AI hasn't escaped gravity: humanoid demos repeat cobot boom failures from a decade ago — yongqianme · 2026-09-17
- Deep dive: what's really happening behind the Tesla Optimus, Figure and 1X NEO humanoid race — MickeySteamboat · 2026-09-17
- Figure teases an 'AI breakthrough' with a demo set for tomorrow — adcock_brett · 2026-09-17
- MessyMem (CoRL 2026): persistent memory for robots via 3D scene graphs and VLM analysis — leto__jean · 2026-09-17
- Fortell founder spent six years building an AI hearing aid startup — RebeccaBellan · 2026-09-17