DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training
Junyan Li, Ruizhi Li, Yu Liu, Xiangshuo Liu, Mingchao Sun, Hongyu Pan, Mu Xu, Lue Fan, Zhaoxiang Zhang
cs.RO, cs.CV
2026-10-09
DreamTrue aligns action renders offline, then reward-tunes counterfactual rollouts, cutting AgiBot interaction defects from 48.12% to 6.25% (nDTW 0.8772).
A robot world model is supposed to work as a simulator. Given the opening multi-view frames, a task instruction, and an action sequence, it has to show the future. Two conditions make that usable. The arm follows the command, and objects move only when contact says they should.
Public logs miss both. Painting an action into the image depends on camera intrinsics and extrinsics. Off calibration puts the rendered arm somewhere other than the arm in the video, so the training target is misaligned before learning starts. Coverage is the other hole. The tapes are mostly successful demonstrations. Missed grasps, slips, and lost support barely show up, and the model lifts an object with the gripper even when the grasp never closed. That hallucination inflates success rates, so action choice and policy scores both come out too high. Collecting a fresh real-robot set with clean calibration and real failures is expensive. DreamTrue repairs the calibration already in the logs and edits the trajectories, without more robot collection.
The system is from NLPR at the Institute of Automation, Chinese Academy of Sciences, and Amap at Alibaba. The generator is a cross-embodiment, multi-view video DiT, a diffusion Transformer. Actions enter through a VACE branch, a side path that feeds visual conditions into the video backbone. Training has three parts: offline calibration, supervised training, and reward post-training.
There is no separate checkerboard session. Recorded joint angles and the initial camera guess render the URDF, the link-and-joint geometry file, into the frame. RoMaV2 finds dense matches, and the segmentation model SAM3 keeps those matches on the arm. The solver updates intrinsics, extrinsics, distortion, and mounting offset. It minimizes reprojection error on the rendered arm, and it requires that two rays seeing the same point, across time or across views, lie in one plane with the line joining the camera centers. The fitted cameras then render the whole action as RGB, depth, and masks. Gripper opening is written into the background color. Camera motion is a Plücker ray map, a dense image of each ray's position and direction. Embodiments do not share an action definition. Once every action is drawn in the same image format, one encoder reads all of them.
Stage I trains on AgiBotWorld-Beta, DROID, RoboMIND 2.0, and RoboTwin 2.0, 2,232 hours and five arm types after filtering. The loss is flow matching: the network regresses the velocity from noise toward the future video.
Stage II adds interactions that were never recorded. One SE(3) perturbation, a rotation plus a translation, moves the final end-effector pose. The arm interpolates from the fixed start, and inverse kinematics emits an action that was not executed. The opening frame stays fixed, so no paired future exists. People label videos from several world models on three axes: a broken arm (L1), a deformed, missing, or duplicated object (L2), and implausible contact (L3), such as an object leaving the table with no grasp. A fine-tuned Qwen predicts a defect probability per axis, and the reward is the negative of that probability. Clips on the original trajectory also receive a PSNR term against the real future, so plausibility cannot be bought by ignoring the pixels. GDPO normalizes advantages inside each sampled group, and DiffusionNFT updates the generator. Released calibration covers 153,666 episodes and more than 1,660 hours on three datasets. The defect set has 44.9K videos, including 30.4K defect labels.
On AgiBot the test set adds 160 counterfactual actions over 54 tasks to held-out recorded actions. Action following is nDTW, similarity after the predicted and commanded trajectories are time-aligned. Three annotators score defects by majority vote, Fleiss' κ 0.7817. Baselines are DreamDojo, GE-Sim 2.0, Genie Envisioner, and EnerVerse-AC. The baseline column is the best value on that metric, so LPIPS is the lower number.
| Metric | Full | No post-training | Best of four baselines |
| Recorded nDTW | 0.8772 | 0.8711 | 0.8124 |
| PSNR / SSIM / LPIPS | 23.06 / 0.936 / 0.094 | 22.66 / 0.928 / 0.099 | 20.72 / 0.890 / 0.112 |
| Counterfactual nDTW | 0.8831 | 0.8735 | 0.7817 |
| Object defect rate | 3.12% | 31.88% | 3.75% |
| Interaction defect rate | 6.25% | 48.12% | 7.50% |
On both recorded and counterfactual actions, the full model leads this set on nDTW and on fidelity. Interaction defects fall from Stage I's 48.12% to 6.25%, 10 of the 160 clips. Object defects fall from 31.88% to 3.12%. Before post-training, 48.12% is still worse than baselines at 16.25% and 7.50%. After it, DreamTrue is just ahead of the strongest of those. The baseline at 7.50% interaction defects reaches only 0.8063 recorded nDTW. EWMScore-P, a pool of WorldArena quality metrics, is 72.51 for the full model, below 72.84 without RL.
On the world-model track of the AgiBot World Challenge 2026, visual quality is 0.6246, action following 0.9651, and the overall score 0.829, first on each. Scene consistency is 0.8974, behind PAIWorld at 0.9041. The overall margin over second place is 0.0045.
Against Ctrl-World on DROID, PSNR moves from 22.00 to 22.95, SSIM from 0.772 to 0.848, and LPIPS from 0.162 to 0.073. RoboMIND 2.0 and RoboTwin 2.0 are reported only as own-model PSNR, 23.84 and 27.35.
The policy check matches the claimed use. Three bimanual RoboTwin 2.0 tasks, easy and hard, four vision-language-action policies (EventVLA, π0.5, X-VLA, starVLA), 240 trajectories and 24 conditions. Predicted success against simulator labels fits y=1.032x-0.009 for DreamTrue and y=0.919x+0.087 for Ctrl-World. The baseline sits above the diagonal wherever success is low, drawing failed grasps as completed. Success-rate MAE drops from 12.92 to 7.08 percentage points. Spearman ρ is 0.937, and the aggregate bias is +0.42 percentage points.
Against dataset-provided calibration, mask IoU rises by 0.228 on average over 130,182 AgiBot episodes (better on 97.1%) and by 0.212 over 63,061 DROID episodes (better on 91.6%). Swapping the new calibration in at inference already helps. Training with it as well adds 1.28 dB of PSNR on AgiBot, to 22.66, and 1.30 dB on DROID, also to 22.66. End-effector sync error falls to 4.60 and 8.06 pixels. Against PointWorld, which needs stereo depth, the RGB-only fit wins IoU on 75.5% of 38,356 DROID episodes. On 4,893 held-out clips the reward model reaches 86.33% accuracy and 74.21 macro-F1. Embodiment macro-F1 is only 69.37, so rare robot defects are not caught as reliably.
Policy scoring with a world model fails when a miss is drawn as a success. Once that upward shift in the low-success region is gone, rankings in simulation sit closer to the simulator labels. Several arms share one checkpoint, and the calibration plus the defect labels are released. The cost that disappears is another real-robot failure-collection campaign. Video-model training is still paid in full.
Interaction defects are 1.25 percentage points under the strongest baseline, and the challenge score leads by 0.0045. Against the previous generation of world models, the gain is modest.
The paper states the limit directly. The action representation, the generator, and the reward all read images. With occlusion or too few views, a contact can look right and still be physically wrong, and the reward can endorse it.
Counterfactual actions only shift the final end-effector pose and interpolate. Slips, deformation, and multi-point contact sit outside that edit. A 6.25% rate on 160 clips moves when a few labels flip. The policy study compares only with Ctrl-World, and only in simulation. There is no real-robot closed loop, and no evidence that a downstream policy trained on these rollouts improves. Cross-embodiment transfer is shown with frames, not with numbers. Object-defect MAE is 0.243, so the reward is still coarse.