IW-ABC trains one policy on 40 tasks to 90.1% mean success

A GPU-Parallel Framework for Heterogeneous Multi-Task Reinforcement Learning

Rui Zhang, Qiwei Wu, Zhengyu Zhang, Tao Li, Hongyu Zhou, Xiang Li, Yunrong Guo, Junjie Lai, Renjing Xu, Weihua Zhang

cs.RO

2026-06-02

Hebero trains one policy on all 40 LIBERO tasks in one GPU-parallel sim. With 50 demos each, IW-ABC reaches 90.1% state-input success, 7.8 points above FAMO-ABC.

What problem this solves

GPU-parallel simulators already run thousands of robot environments on one accelerator. Most of that throughput still trains one specialist on one task. A single policy that must acquire many differently structured skills hits three tighter constraints: sparse success signals, tasks that learn at uneven speeds, and a small demonstration set that could initialize the policy, regularize it, or define the reward.

RLBench and LIBERO offer diverse, demonstration-rich tasks, but their physics runs on CPU and cannot host thousands of heterogeneous scenes in one process. MTBench moves Meta-World onto the GPU, mainly for state-based multi-task RL. ManiSkill3 has heterogeneous GPU simulation and visual RL, with reference learners trained per task. Hebero, a heterogeneous robot-learning benchmark, puts all 40 LIBERO tasks, state and visual inputs, demonstrations, and one evaluation protocol into a single Isaac Lab process.

Evaluation also drops LIBERO's fixed starts. Each movable object's planar position is redrawn uniformly inside the box spanned by the 50 demonstration starts, then rejected on collision. The test is placement robustness, not replay of stored trajectories.

Method

The benchmark and the learner are separate. Hebero fixes the tasks, the simulator, and the evaluation contract. DGPO (Demonstration-Guided Policy Optimization) is one training stack on that contract. Demonstrations supply a dense tracking reward and a privileged critic. The backbone is task-conditioned PPO. Learners differ only in how they consume the demonstrations.

The 40 tasks are ten each from LIBERO's Goal, Object, Spatial, and Long suites. Scenes and success conditions differ. The action space, control rate, and robot body do not. The task code is a 36-dimensional multi-hot vector: six shared subtask bits and 30 task-ID bits. The action has seven dimensions, six for relative end-effector pose and one for the gripper. State actors also receive a 26-slot object-target pose buffer, 300 inputs in total. Visual actors replace that buffer with pooled features from a frozen ViT on third-person and wrist images, 450 inputs. The actor is a [512, 256, 128] MLP, about 0.32M parameters in the state setting.

The task reward pays only on success and penalizes jerky actions and joint-limit violations. The tracking reward adds an exponential kernel on end-effector, joint, gripper, object-pose, and articulation error against the demonstration, so both the motion and its effect on the scene get credit. The critic also sees those errors and the demonstration cursor. Quantities that deployment cannot observe stay in the value network. That is asymmetric value learning.

IW-ABC drives imitation and task balance from one cheap signal. For each task, tauk is an exponential moving average of success on training episodes, with no separate evaluation pass. ABC (adaptive behavior cloning) lowers the imitation weight as tauk rises and as global training time advances. A task solved early sheds the demonstration before the global schedule ends. After annealing, even a hard task keeps only a residual floor, so the policy is not pinned to expert actions indefinitely. The target is the action at the current demonstration cursor. The loss is a weighted squared error on the Gaussian mean, added to PPO.

IW (importance weighting) steers the PPO policy, value, and entropy terms toward lagging tasks. Rollouts stay evenly split across tasks. When tauk sits below the multi-task mean, a sigmoid raises that task's weight into roughly 0.5 to 2. One backward pass is enough. There are no per-task gradients. FAMO-ABC, the strongest baseline, adapts weights from loss progress and applies them only to the critic, while imitation stays ABC, under the same budget. The same stack also includes BC warmup, a DAPG likelihood regularizer, and an RFCL-style reverse reset curriculum.

Results

On one L20 with 1,600 environments, Hebero simulates at about 7,500 steps/s and 10.6 GiB. The 96-core MuJoCo baseline, at the same task and environment count, does 2,200 steps/s and about 1.5 TiB of host memory. GPU throughput is about 3.4 times that figure. Eight L20s and 25,600 environments reach 78,500 steps/s end to end. On one GPU, more replicas per task raise mean success inside a fixed wall-clock budget.

Every learner shares DGPO: 50 demonstrations per task, 30,000 online PPO iterations, three seeds, eight L20s with 2,000 environments each, about two days.

MethodMean SRLong SRCoverage (≥80%)SR-AUC
PPO50.8±3.1%10.0%20/4045.2%
BC→PPO47.5±2.6%20.0%18/4040.8%
ABC75.5±3.6%40.0%29/4071.9%
FAMO-ABC82.3±2.9%53.3±5.8%30/4075.6%
IW-PPO44.9±2.5%20.0%18/4041.3%
IW-DAPG68.5±3.0%30.0%27/4062.4%
IW-RFCL66.4±2.8%30.0%26/4060.1%
IW-ABC90.1±3.8%70.0±10.0%35/4082.1%
Vis IW-ABC93.5±2.6%81.9±11.5%38/4083.3%

Mean SR averages all 40 tasks. Long SR covers the ten long-horizon tasks. Coverage counts tasks at or above 80% success. SR-AUC is the normalized area under Mean SR across PPO iterations.

State-input IW-ABC reaches 90.1%, 7.8 points above FAMO-ABC at 82.3%. Long SR moves from 53.3% to 70.0%. Among the ten Long tasks, those at or above 80% number 4 for ABC, 7 for IW-ABC, and 9 for the visual policy. The visual policy, frozen encoder included, has 5.92M deployed parameters, 93.5% mean success, and 38/40 coverage. Weighting only the actor, not the critic, drops the same recipe to 78.2% mean success, 50.0% Long SR, and 31/40 coverage.

BC warmup followed by PPO scores 47.5%, below plain PPO at 50.8%. IW without ABC drops mean success to 44.9% while lifting Long SR from 10% to 20%. The two pieces have to move together. With IW held fixed, ABC still beats DAPG at 68.5% and the RFCL curriculum at 66.4%.

The physical test is a different four-task set on a Piper arm, taken from RoboTwin. One jointly trained state policy is deployed with frozen weights. Object pose comes from an external FoundationPose pipeline. Over 20 trials per task the counts are 20/20 for clicking a bell, 15/20 for shaking a bottle, 13/20 for moving a pill bottle onto a pad, and 18/20 for placing a container on a plate. Overall success is 82.5%.

Why it matters

Hebaro is a shared contract: heterogeneous tasks, GPU-parallel rollouts, and one multi-task protocol, with the demonstration interface left swappable. Reward, observations, PPO, and evaluation stay put, so learner differences are readable. IW-ABC adds little machinery. The progress signal is a moving average of training success. A 0.32M state actor at 90.1% says network capacity is not the first constraint on these 40 tabletop skills.

The gain is incremental, and it sits on tracking rewards plus online imitation. Importance weighting alone lowers mean success. The visual 93.5% uses a privileged critic and a frozen encoder, so it is not an end-to-end vision result. The 82.5% robot number shows that four state-based tasks with external pose estimation can leave simulation. It does not show that one visual policy already covers dozens of physical skills.

Limitations

The paper flags two limits. Time-indexed demonstration actions fall out of alignment once the policy leaves the reference trajectory, and the tracking reward only reduces that mismatch. Compute also confines training and evaluation to LIBERO and RoboTwin. Harder and larger task sets are left for later.

A few readings should stay narrower than the headline. The 40 evaluation tasks are the training tasks. There is no held-out set, so the numbers measure multi-task mastery, not transfer. FAMO reweights only the critic. The policy loss and entropy stay uniform, and the 7.8-point gap belongs to that pairing. Long SR is noisy: ±10.0 points for IW-ABC and ±11.5 for the visual run, across three seeds, which is too thin to lock in a long-horizon ranking. Real-robot failures are not split between the policy and the external perception stack. Single-GPU scaling is a trend in a figure, without a success rate at each environment count.

Terms

Source

Related papers

All paper explainers