DeepLeap's DELE-w0.5 ditches video-generation pipelines for robot manipulation

jiqizhixin · x · 2026-09-13

DeepLeap releases DELE-w0.5, arguing that video-generation-based world-action models waste enormous compute predicting high-dimensional visuals just to output low-dimensional actions. DELE-w0.5 abandons that pipeline entirely: it jointly learns robot actions without future visual trajectory prediction, instead learning goal-conditioned behavior reorganization — preserving the task goal and adapting actions when the environment deviates. The architecture flows from language instruction plus observation to VLM task planning to action execution.

Original post →

More from Embodied

Embodied channel →