Magic-W0: a structured world-action foundation model topping RoboDojo-Sim at 27.10

Xuhua Chen · hf · 2026-10-07

Magic-W0 is a world-action foundation model that jointly models structured physical state evolution and continuous actions. Interaction is represented as a Structured World Transition—Current State (vision-language context plus current 3D geometry), Transition (3D Motion capturing action-induced changes), and Future State (task-relevant Future Semantics).

A layer-aligned world-action architecture couples prediction and control: evolving action hypotheses condition world-transition prediction, while predicted world representations inform action generation. It is pre-trained on large-scale egocentric human manipulation, UMI, real-robot, and simulation data, with latent supervision from pre-trained visual models.

Inference-time interventions show structured world representations respond systematically to candidate-action changes. Magic-W0 scores 27.10 on RoboDojo-Sim (highest among compared WAMs) and shows strong real-robot performance after fine-tuning with limited downstream data.

Original post →

More from Embodied

Embodied channel →