Hard real-robot tasks: SimpleICL 68% vs Fast-WAM 42% and pi0.5 31%

In-context Robot Learning Made Simple: A Democratized Recipe for Manipulation Tasks

Minxing Li, Minghao Han, Weizhi Zhao, Hanwen Wang, Xiangshuo Liu, Shuyao Shang, Jingxiang Zhou, Mingchao Sun, Hongyu Pan, Mu Xu, Yu Liu, Lue Fan, Zhaoxiang Zhang

cs.RO

2026-09-30

SimpleICL names four cues to copy from human video, trained with cheap cross-group pairs. Hard success on eight real tasks is 68%, vs 42% Fast-WAM and 31% pi0.5.

What problem this solves

Watch a person, then do the task, with no new teleop set and no finetune. That is the bet in robot in-context learning. HOST, GEN 1.5, and Zero-WAM already report early zero-shot behavior. A demonstration video still mixes the hand path, object identity, grasp site, action order, and the end state.

Shift the object and a copied path misses. Optimize only the end state and the order drops out, as does the choice between a mug handle and the mug body. The language in Figure 1 is one line, Put the tea bag into the cup. The missing piece is a boundary: what the policy must follow, and what it must drop.

Method

SimpleICL comes from the Institute of Automation, Chinese Academy of Sciences, and Amap at Alibaba.

Four factors are in bounds. Action is a category: grasp, place, rotate, press, or translate, not the joint curve. The object is handled at its pose during execution, not at the absolute coordinates of the demo. Order is part of the task. Banana then apple is a different task from the reverse. Affordance is whether the grasp is on the cup body or the handle.

Three factors are out of bounds. Ordinary position shifts between demo and execution are allowed. Extreme shifts count as a different meaning, and the paper gives no distance cutoff. Human speed, arm shape, appearance, grasp angle, and reaching habits are ignored. So are lighting, background, and clutter.

The data is built on that boundary. One scene includes both pressing a soap dispenser and picking it up onto a tray. Object positions are perturbed, other objects sit nearby, and each scene contains a single target instance. The same action set is collected in different orders. Affordance follows a human-robot consistency rule: within a pair the hand and the gripper grasp the same spot, and the next pair switches to the handle or the body.

Invariance is augmentation. Demonstrators span bare hands, colored gloves, and different skin tones, and one robot trajectory is paired with videos of different light, color temperature, and camera pose. Collection is grouped. A group is one scene plus the tasks in it. A major switch replaces the scene. A minor switch only changes layout, lighting, or distractors, and human videos are then re-paired with robot trajectories across groups. One platform produces up to about 3,000 valid pairs in a working day.

The backbone is Wan2.2-5B. Following Fast-WAM, a 1B action DiT is interpolated in with horizon 32 and connected to the video DiT through MoT. A module of 128 query tokens and 4 cross-attention blocks, hidden size 3072, about 0.6B parameters, compresses the demo to a fixed length and feeds both contexts. Those contexts already contain text and state. The video DiT is trained with LoRA. Everything else is fully trained. Camera views are concatenated into one image.

Results

Simulation runs in RoboTwin 2.0. Native layouts often leave unrelated objects outside the workspace. The ICL-oriented set has 18,000 demos, against 27,500 native demos. Mean success on semantic discrimination: Fast-WAM on native data 37.6%, SimpleICL on native data 40.3%, SimpleICL on ICL-oriented data 72.1%. Removing video conditioning leaves 61.6%. Removing future prediction drops the score to 20.2%.

On seven unseen tasks the means are Fast-WAM 24.6%, SimpleICL 65.9%, no video 26.3%, and no future prediction 17.9%. Without video, new tasks fall back near the baseline. Without future prediction both evaluations collapse. In simulation the demonstrator and the executor are both AGILE X, so there is no jump from a human hand to a gripper.

The real platform is a Franka with three RealSense D435 cameras, teleoperated with a SpaceMouse. The eight tasks are outside the training set, and so are the sponge, plate, drawer, fruit, bowls, shelf items, soap dispenser, and target cup. Easy is one task with one legal execution. Medium puts several tasks in the same view. Hard gives one task several legal modes, for example two adjacent plates of which only the demonstrated one should be wiped. Table 3 stores success as a decimal. The percentages below use that same scale.

SettingSimpleICLFast-WAMπ0.5
Easy mean81%80%72%
Medium mean70%65%47%
Hard mean68%42%31%

The Easy mean is 1 point above Fast-WAM. Wiping a plate scores 77%, below π0.5 at 97% and Fast-WAM at 87%. Serving tea scores 73%, below 83% and 80%. On Hard, SimpleICL leads every task: chopsticks 77% versus 27% and 10%, tea 57% versus 27% and 20%.

Intent-following averages 89% with discriminative collection and 75% without it. Composition barely uses that design. Snacks, cups, and bags stay at 93%, 97%, and 93% when the extra orders are removed, against 97%, 100%, and 90% when they are kept. Affordance moves a lot: handled cup 87% versus 60%, tape measure 77% versus 40%, tape 83% versus 47%. Action sits in the middle. Tissue falls from 83% to 67%. Chopsticks do not, at 93% versus 97%.

The full model scores 81% in the original setup, 79% if the object is shifted, 77% with a different hand appearance, 80% with another person and motion style, 78% under new lighting, 79% with a new background, and 75% with clutter. Without cross-group pairing the original setup is still 80%, a shift falls to 51%, clutter to 66%, and background to 71%. The policy then reaches for the coordinates shown in the human video.

Why it matters

The reusable piece is grouped collection plus cross-group pairing, and the data and training pipeline are slated for release. When a scene contains a single task, Fast-WAM is already at 80%, and SimpleICL adds 1 point. When several tasks share the view, Hard opens a gap: 68% versus 42% and 31%. Short language stays in the context. The video supplies object, order, and grasp. The network change is incremental: Wan2.2-5B, plus a 1B action DiT and a 0.6B query module.

Limitations

HOST, GEN 1.5, and Zero-WAM are not in the comparison. Part of the Hard lead over Fast-WAM and π0.5 is that those baselines do not receive a video. The paper also does not state which real-robot data those baselines were trained on. Both sides of the simulation use AGILE X. The human-to-gripper gap is measured on one Franka and eight tasks. Success rates have no error bars, and the number of trials per cell is unreported.

Training assumes a single instance of the target in each scene. Two identical cups are outside the definition. An extreme shift is called a new meaning, with no measured threshold. Speed and a specific approach angle are discarded on purpose. Affordance does not emerge on its own. The same object needs both grasps in the data, or intent-following falls into the 40% to 60% band. Human-robot consistency further forces the hand and the gripper to grasp the same point in every pair. On Easy, plate wiping and tea serving lose to at least one baseline, which conflicts with the claim in Figure 1 that all eight tasks are won.

Terms

Source

Related papers

All paper explainers