Workhorse trains G1 from retargeting-free human demos to kick boxes, catch throws, climb a suitcase

Workhorse: Learning Robust Whole-Body Humanoid Loco-Manipulation from Human Data

Songbo Hu, Qiayuan Liao, Yufeng Chi, Kevin Zakka, Yakun Sophia Shao, Pieter Abbeel, Koushil Sreenath

cs.RO, cs.LG

2026-10-07

Human demos recorded with wearables, no retargeting, train a visual planner and an RL tracker that drive a real Unitree G1 to kick boxes, catch throws, climb a suitcase; 77% sim success.

What problem this solves

Humanoids can already track a given whole-body reference motion well; BeyondMimic and GMT are built for exactly that. Manipulation tasks come with no reference. From egocentric RGB and proprioception alone, the robot has to plan its own whole-body motion: squat, reach, brace a leg against an object, and stay up when shoved.

Data is the other half of the problem. Teleoperation ties up a robot and a skilled operator for every demonstration. Robot-free wearable rigs (HuMI and successors) are far cheaper, but in the tasks published so far the feet only walk, and the motions are mostly quasi-static. This paper puts the feet to work: kicking a box into a shelf, toppling a suitcase with a leg, stepping onto it to reach higher.

Method

Two policies talk through one interface: the poses of five links, meaning the torso, both wrists, and both feet.

Capture hardware is light. Five HTC Vive trackers sit on the chest camera mount, both hands, and both shoes; four Lighthouse base stations localize them at 60 Hz in one gravity-aligned frame over a 5 m × 5 m area. A chest-mounted DJI Osmo Action 4 records 129° near-fisheye video, and the same camera model goes on the robot, so training and deployment share the same distortion and rolling shutter. The protocol perturbs objects only, never the demonstrator, so every recorded pose is intended motion. Three claps align the two clocks to within 5.5 ms. One demonstrator recorded everything: 2.2 h for box sorting, 1.2 h for the suitcase task, 13 min for catching.

The interface holds no joint angles. The recorded human motion itself is the training target for both policies: no retargeting, no scaling, and the robot must hit the exact poses the human hit. The only robot-specific quantities are five constant "mounts" that pin the alignment frames onto the robot's links. Porting to a new humanoid means measuring those five constants, swapping the robot model, and retraining from the same demonstrations.

The visual planner reads four images and nine five-link history samples, and outputs a 58-step, 1.16 s action chunk; it is a flow-matching policy with a DiT backbone, a ResNet-18 image encoder, 10 Euler steps, replanning at 5 Hz. The whole-body tracker is an RL policy built on BeyondMimic, running at 50 Hz on proprioception alone. Its reward scores the wrists in the world frame, so they must arrive, and the feet in the torso frame, so the robot is free to step wherever balance needs.

Because the two policies train separately but must cooperate in closed loop, each one's data augmentation rehearses the other's deployment errors. Offline, SAM 3 cuts the person out of the recorded images, LaMa inpaints, MoGe-2 estimates metric depth, and MuJoCo renders the robot into the scene; at training time the planner's five-link history gets a random-walk drift, with images warped consistently through the depth map. The tracker trains on commands that drift and jump: replans arrive at random 0.2-1 s intervals, drift grows to at most 10 cm and 10° of yaw between replans, and the reward is computed against the drifted command. At deployment, timestamps align the two policies' delays, and the overlap between consecutive chunks is inpainted during flow integration.

Results

On a real Unitree G1, with the tracker onboard and no external motion capture, the robot places two boxes on a shelf by hand and kicks the third into the bottom shelf, catches a thrown box with about 0.56 s of flight time, pushes a 14 kg rolling suitcase, topples it with a leg against its base, steps up 0.39 m onto it, grabs a box, and jumps down. A person pushes, kicks, and takes boxes away mid-task; the robot stays up and continues.

Success rates are simulation-only: a Gaussian-splat reconstruction of the demonstration room, 100 episodes per setting, with 40 N·s impulses in random directions every 5 to 10 s.

VariantNo pushesWith pushes
Full system77%64%
Joint command (online IK)54%42%
Link-local action frame5%4%
No planner augmentation20%0%
No command augmentation76%52%

A few readings. Swapping the five-link command for an online-IK joint command (the BifrostUMI approach) costs 23 points of success, yet cuts falls under pushes from 15% to 4%: a joint command is always reachable, so the robot falls less but loses the task more. Without projected gravity, success drops to 1% and the robot falls in almost every episode. The command augmentation barely matters without pushes (76% vs 77%) and earns its keep under them (52% vs 64%). Retrained from the same demonstrations, a simulated Unitree H2, a full-size humanoid closer to the demonstrator's proportions, reaches 83% without pushes against the G1's 77%; the gap is smaller than the ±11-point 95% interval for such a difference.

Why it matters

Teaching dynamic, feet-involved manipulation from robot-free human data sidesteps the cost of teleoperation, and the five-link interface plus fixed mounts turns cross-embodiment transfer into measuring five constants and retraining; the H2 result is early evidence that this works. The augmentation recipe, each policy training on the errors the other will make, applies to any planner-plus-tracker stack, humanoid or not.

Keep the position in view: these are three task-specific policy pairs, not one general model, and every success-rate number comes from a simulation of a single room.

Limitations

The authors list them plainly. Each task has its own planner and tracker. One demonstrator recorded all data in one room, and all tests happen there or in its reconstruction. Real-robot experiments demonstrate capability; rates come from simulation only. The H2 result is simulation-only, with the real robot as future work. The five links carry no fingers and no head, so the interface cannot express a grasp, and the camera sits fixed on the torso with no neck to look around. Lighthouse base stations confine recording to one room.

Reading closely adds more. In the simulated evaluation, appearance, lighting, friction, and mass stay fixed across episodes; only box placement and start pose vary, so generalization beyond the room is untested. About 15% of the planner's training data had to come from the reconstructed room to close the appearance gap, which says the sim images still differ from the real ones. A 15% fall rate under pushes is far from deployable. Contact-heavy pieces like the kick and the suitcase topple depend on how well the objects are modeled, and the paper does not validate that separately.

Terms

Source

What people are saying

Related papers

All paper explainers