WorldDiT unifies robot action generation and future frame prediction in one diffusion model
bageldotcom · hf · 2026-07-28
WorldDiT proposes a unified diffusion transformer for robot world and action modeling
Bagel Labs introduces WorldDiT, a single diffusion transformer that jointly learns continuous action chunks and future visual patches, instead of relying on a large pretrained VLM as the action backbone.
Key results:
- Evaluated on four LIBERO simulation suites
- Reported to sit on the Pareto frontier for total parameters vs. mean success among methods that report all four suites
- Positioned as a sub-billion-parameter baseline for future scaling studies
More from Embodied
- NVIDIA points Jetson users to a guide for running GenAI locally on-device — NVIDIAAI · 2026-07-28
- Meta upgrades Ray-Ban Display glasses with Muse Spark AI and Threads integration — emmanuelvivier · 2026-07-28
- Meta Ray-Ban Glasses Updated with Threads Integration and Muse Spark AI — emmanuelvivier · 2026-07-28
- NVIDIA says Cosmos models have passed 10 million downloads on Hugging Face — nvidia · 2026-07-28
- NVIDIA says Jetson now fits in a bag while powering robots and edge AI — nordicinst · 2026-07-28
- Austin is becoming a major robotics hub with UT Austin, Tesla and Apptronik — lukas_m_ziegler · 2026-07-28