ModAR predicts robot futures one modality at a time; 30M model beats 6B video model

k7agar · x · 2026-09-17

World-action models usually imagine the future in RGB, but pixels spend capacity on details irrelevant to robot policies. ModAR instead predicts the future one modality at a time — DINO features, point tracks, depth — with each prediction informing the next. Trained from scratch, its 30.1M-parameter model outperforms a 6B video-model-initialized baseline finetuned on the task.

Related event: ModAR challenges RGB representations, beating 6B video models with 30M parameters(3 posts)→

Original post →

More from Embodied

Embodied channel →