Deepleap's DELE-w0.5 world-action model hits 62.5% real-robot success, 2.1x best baseline, without video generation

机器之心 · wechat · 2026-09-04

Deepleap (深度跃迁) released DELE-w0.5, a world-action model demonstrating "goal-conditioned behavior reorganization": when a microwave door is closed mid-task, the robot reopens it; when it's left ajar, the robot pushes it open with the popcorn bag — preserving task goals under off-training-distribution perturbations.

Core idea: video-based world-action models jointly predict actions and future visual trajectories, where high-dimensional intermediate frames waste compute and throttle control loops. DELE-w0.5 instead jointly predicts the future world representation after action completion — a DINO-v3 latent encoding object positions and spatial relations, with no video generation. During training this branch supervises the policy toward outcome-oriented state transitions; at inference it's removed entirely, yielding a 87.5 ms median latency on an RTX 4090.

Real-robot results (Astribot S1, 4 long-horizon tasks × 8 methods × 20 trials each, 640 experiments):

The model base is built from scratch (no repackaged open-source VLA/video weights). The company, focused on cross-embodiment embodied brains, recently closed a tens-of-millions-RMB angel round led by Fuxing Chuangfu.

Original post →

More from Embodied

Embodied channel →