Magic-W0: a structured world-action foundation model topping RoboDojo-Sim at 27.10
Xuhua Chen · hf · 2026-10-07
Magic-W0 is a world-action foundation model that jointly models structured physical state evolution and continuous actions. Interaction is represented as a Structured World Transition—Current State (vision-language context plus current 3D geometry), Transition (3D Motion capturing action-induced changes), and Future State (task-relevant Future Semantics).
A layer-aligned world-action architecture couples prediction and control: evolving action hypotheses condition world-transition prediction, while predicted world representations inform action generation. It is pre-trained on large-scale egocentric human manipulation, UMI, real-robot, and simulation data, with latent supervision from pre-trained visual models.
Inference-time interventions show structured world representations respond systematically to candidate-action changes. Magic-W0 scores 27.10 on RoboDojo-Sim (highest among compared WAMs) and shows strong real-robot performance after fine-tuning with limited downstream data.
More from Embodied
- 2023's ACT + ALOHA was the original physical AI era, says dev — lukas_m_ziegler · 2026-10-07
- EAPN: execution-aligned noise fixes mode switching in asynchronous replanning, 96.7% bimanual success — Di Wu · 2026-10-07
- Two-step flow denoising cuts VLA inference from 61.6ms to 22ms for real-time robot control — Di Wu · 2026-10-07
- CTP: single-pass multimodal robot policy hits 97.25% on LIBERO, cuts inference latency to 75.8ms — Di Wu · 2026-10-07
- RaiSim is back: v2.8.0 brings ultra-fast accurate CPU simulation plus photorealistic rendering — ChongZzZhang · 2026-10-07
- LiteReality-Agent turns images into simulation-ready 3D scenes for MuJoCo robot planning — elliottszwu · 2026-10-07