CrossBFM Distills a Shared Behavior Space Across Humanoid Robots in Under One GPU-Hour

Tan-Dzung Do · hf · 2026-09-30

CrossBFM treats the latent behavior space of Behavior Foundation Models as the transferable asset across humanoid embodiments. Using retargeting for frame-level cross-embodiment correspondence, a unified robot-agnostic encoder distills the space to multiple robots in under one GPU-hour, with latent-conditioned PPO trackers adding 10 GPU-hours. All three prompting modes—motion tracking, goal reaching, reward optimization—transfer across three humanoids; training on a subset of robots recovers up to 89% performance on unseen ones, verified on real hardware.

Original post →

More from Embodied

Embodied channel →