Alibaba's RynnValue robot value model uses temporal distance to boost real-world success to 72.5%

Alibaba-DAMO-Academy · hf · 2026-08-11

Alibaba DAMO Academy introduces RynnValue, an open-source value foundation model for robotic manipulation that replaces task-internal anchors with temporal distance (directed cost-to-go to language-specified goal). Labels derived from timestamps enable scaling to 7,000+ hours and 3M instruction-conditioned clips without preference annotations. Combining random temporal sampling, temporal-order shuffling, and value-isolation attention, it achieves average Kendall's taua of 0.675 on RBM-EVAL-OOD, surpassing preference-supervised SOTA (0.655) and more than doubling progress-only (0.292), with zero-shot generalization. Converted to dense rewards, it raises real-world policy success from 52.5% to 72.5% online and 63.8% to 82.5% offline.

Original post →

More from Embodied

Embodied channel →