DeepLeap's DELE-w0.5 ditches video-generation pipelines for robot manipulation
jiqizhixin · x · 2026-09-13
DeepLeap releases DELE-w0.5, arguing that video-generation-based world-action models waste enormous compute predicting high-dimensional visuals just to output low-dimensional actions. DELE-w0.5 abandons that pipeline entirely: it jointly learns robot actions without future visual trajectory prediction, instead learning goal-conditioned behavior reorganization — preserving the task goal and adapting actions when the environment deviates. The architecture flows from language instruction plus observation to VLM task planning to action execution.
More from Embodied
- XPeng's humanoid robot IRON leaves the assembly line, mass production targeted before end of 2026 — emmanuelvivier · 2026-09-13
- Mecka AI nears $500M valuation amid race for robot training motion data — emmanuelvivier · 2026-09-13
- The Smoothest Robot Rollout of 127 Episodes Still Failed: Motion Smoothness Metrics Measure Motion, Not Success — Scobleizer · 2026-09-13
- freecad-mcp: 2.2k-star MCP server lets Claude Desktop drive FreeCAD for CAD and FEM — tom_doerr · 2026-09-13
- Robots roll out: AI patrol bots on Chinese streets, Sally teaching assistant in US schools — Scobleizer · 2026-09-13
- UBTECH scales humanoid production toward 10,000-unit capacity as U1 orders top 13k — CyberRobooo · 2026-09-13