ModAR predicts robot futures one modality at a time; 30M model beats 6B video model
k7agar · x · 2026-09-17
World-action models usually imagine the future in RGB, but pixels spend capacity on details irrelevant to robot policies. ModAR instead predicts the future one modality at a time — DINO features, point tracks, depth — with each prediction informing the next. Trained from scratch, its 30.1M-parameter model outperforms a 6B video-model-initialized baseline finetuned on the task.
More from Embodied
- ActionPiece rethinks action tokenization for VLA models, hits 94.8% on LIBERO — DeepCybo · 2026-09-17
- Tesla Optimus passes CAPTCHA designed to prove you're not a robot — TansuYegen · 2026-09-17
- Bionic robotic fish flexes body like a real fish at China trade fair — rohanpaul_ai · 2026-09-17
- GPT-6 Astra reportedly pretrained on 100k+ GPUs at Stargate, with big real-to-sim implications — erwincoumans · 2026-09-17
- UBTECH launches hyper-realistic U1 companion humanoid robots, sparking privacy questions — TansuYegen · 2026-09-17
- The egocentric data boom: Maxinsights has delivered 2M+ hours for robot training — 机器之心 · 2026-09-17