WorldDiT unifies robot action generation and future frame prediction in one diffusion model

bageldotcom · hf · 2026-07-28

WorldDiT proposes a unified diffusion transformer for robot world and action modeling

Bagel Labs introduces WorldDiT, a single diffusion transformer that jointly learns continuous action chunks and future visual patches, instead of relying on a large pretrained VLM as the action backbone.

Key results:

Original post →

More from Embodied

Embodied channel →