No new demos: causal traces lift π0.5 real-robot success by 82.2 points

ARC: A Reasoning Recipe for Robot Foundation Models

Gokul Puthumanaillam, Tao Sun, Elie Aljalbout, Moritz Reuss, Zhaoshuo Li, Fabio Ramos, Ankit Goyal, Jenai Xuning Yang

cs.RO, cs.AI

2026-10-09

ARC relabels DROID with action-grounded causal traces, then fine-tunes π0.5 and Cosmos3-Nano. Reasoning success rises up to 50 points; real-robot π0.5 jumps from 9.5% to 91.7%.

What problem this solves

Robot foundation models still improve by the expensive route: more demonstrations, larger networks, another foundation-scale training run. A vision-language-action model such as π0.5 turns images and an instruction into motor commands. A world action model such as Cosmos3-Nano-Policy also predicts how the scene will look next. Manipulation already works. Another point of success still has to be bought with data and training.

ARC takes a side path. The architecture stays put, no new robot trajectories are collected, and pretraining is not repeated. The claim is that the right reasoning supervision can lift a model that already exists. The wrong text does not. Plans, object lists, and subtask labels are often predictable from the same image and the same instruction, so a language loss can be satisfied while the action expert keeps reading vision. In the paper's attention check, relative attention from π0.5's action expert onto language is 4% for subtask labels and 6% for Embodied CoT, the format that writes a plan, objects, and a motion direction as intermediate text. A trace that actually steers control has to name the next action, say why it fits, and say what should become true once it is done.

Method

The recipe has three parts: the content of the trace, the source of labels, and the fine-tune that makes a pretrained controller use them.

Each trace is natural language tied to one action chunk. Seven fields, State, Cause, Consequence, Effect, Action, Avoid, and Completion, record the relevant facts, why they matter, a look ahead, the change this step should cause, the motion that produces it, what must not happen, and whether the task is finished. The policy reads that bundle as one paragraph of about 700 characters.

DROID stores observations and actions, not reasons. The v1.0.1 training split has 95,658 episodes. Gemini 2.5 Pro watches each video, writes a time-aligned narrative, and picks keyframes, so a frame-level label knows how far the task has already gone. GPT-5.5, at low reasoning effort, writes episode-specific state checks and then a trace aligned to the action chunk. The run covers about 74,740 episodes (78.1% of the split) and about 1.2 million frames, roughly 16.1 annotations per episode: ARC-Trace-DROID. Those checks are model-written descriptions, not measured state.

Fine-tuning keeps flow matching. The model predicts a velocity that moves a noisy action back onto the demonstration; the world action model also moves future observations back. π0.5-DROID uses PaliGemma, SigLIP plus Gemma 2B, and an action expert of about 300M parameters. A two-layer encoder appends the trace to the image-and-instruction context. The vision encoder, language backbone, and action expert are updated together. Factual traces alone barely change the action when the text contradicts the scene, so about 5% of traces are swapped for counterfactuals that oppose the demonstrated action. Those samples leave the prediction loss and only push a one-step action estimate away from the demonstration. Cosmos3's 8B reasoner already changes its action distribution under conflicting text, so that penalty is dropped. Training keeps the prediction loss plus a next-token loss on the trace. A window is one observed frame plus 32 future frames.

At test time the fine-tuned weights stay frozen. An external vision-language model writes the trace from the instruction and the live cameras. The default writer is Qwen3.6-35B-A3B. Updates are asynchronous, and a rejected trace leaves the previous one in place. A near-real-time π0.5 loop adds about 10 ms. In the hardware example the trace refreshes near 1 Hz while control runs at 15 Hz.

Results

Every run uses the DROID embodiment. Fine-tuning sees only ARC-Trace-DROID. The three simulation suites and the real robot get no further training.

On RoboLab-120 the same tasks are phrased as vague, default, or specific instructions. The two ARC models take the top two places in all three bands.

InstructionModelSuccessComparison
Vagueπ0.5+ARC45.1%base 15.2%, +29.8 points
VagueCosmos3-Nano+ARC44.9%base 20.6%, +24.3 points
DefaultCosmos3-Nano+ARC48.8%base 36.8%, +12.0 points
Defaultπ0.5+ARC45.3%base 28.0%, +17.3 points
SpecificCosmos3-Nano+ARC51.2%base 39.7%, +11.5 points
Specificπ0.5+ARC45.0%base 28.1%, +16.9 points

Base π0.5 falls from 28.1% on specific instructions to 15.2% on vague ones. With ARC it stays near 45% in every band, and on vague instructions π0.5+ARC ranks first at 45.1%.

On MolmoSpaces, Cosmos3-Nano+ARC reaches 57.6% (+18.6 points) and π0.5+ARC reaches 45.0% (+27.2 points over the 17.8% DROID base). Both leaderboards are reported as new best results.

RoboLab-Reasoning-50 is built in this paper. It covers contextual understanding, long-horizon memory, discovery by interaction, and negation.

MethodSuccess
Cosmos3-Nano+ARC61.0% (base 11.0%, +50.0 points)
π0.5+ARC57.0% (base 10.2%, +46.8 points)
π0.5 + explicit subtasks21.0%
Cosmos3-Nano-Policy11.0%
π0.510.2%

Explicit subtasks are language supervision too. On the same π0.5 backbone they reach 21.0%. Causal traces reach 57.0%.

Hardware rebuilds 28 of those tasks with no extra fine-tuning. π0.5+ARC scores 91.7%, against 9.5% for π0.5 (+82.2 points) and 11.9% for Cosmos3-Nano-Policy. Cosmos3+ARC is not reported on the robot.

On default RoboLab-120, π0.5+ARC at one Euler step scores 27.6%, 0.4 points under its 10-step base of 28.0%. Cosmos3-Nano+ARC at two UniPC steps scores 41.6%, 4.8 points above its 4-step base of 36.8%, then drops to 18.0% at one step. On the open Cosmos3 training stack, ARC reaches the instruction-only baseline's 10k-update success with about 4.3× fewer updates.

The ablations match that picture. The causal fields say which state change matters; the action field ties that change to a physical step. Drop either side and the controller cannot fill the gap. Updating only the action head, or only the generator, lags a joint update of the language backbone and the controller. Swapping the external trace writer, with π0.5+ARC held frozen, leaves success in the same band. The result is not tied to the default Qwen3.6.

Why it matters

A lab that already has a VLA or a WAM does not have to collect another robot dataset, or repeat foundation-scale training, just to add this kind of reasoning. The video and the actions are already in the old demonstrations. The trace is ordinary text, so generation can use a standard language-model stack, and the control loop can stay at 15 Hz.

Instructions can stay coarse. The trace states the change the next action should cause, which is why success barely moves from vague wording to specific wording. Remembering progress, retargeting when the scene changes, and retrying a failed grasp are the same paragraph being refreshed during execution.

The gains are uneven. Cosmos3 picks up 11.5 points when the instruction is already specific, and the reasoning suite is where the lift reaches 50 points. This is a recipe for an intermediate representation. The controller itself is the same generation as before.

Limitations

The trace stays outside the policy. The built-in language stack can follow a trace and cannot reliably write one, so deployment keeps an external vision-language model in the loop. The conclusion names the follow-ups: policies that generate their own causal traces, and robots other than the one tested here. Simulation and hardware are both DROID only.

The 50-point jump is on Reasoning-50, a suite introduced in this paper. RoboLab-120 and MolmoSpaces also move, but on specific instructions Cosmos3 gains 11.5 to 12.0 points. Hardware at 91.7% against 57.0% for the same π0.5+ARC model in simulation is awkward to read: both bases sit near 10%, and the adapted policy is what splits. The 28 hardware tasks are the ones that could be physically rebuilt. The paper does not show that this subset matches the full suite in difficulty.

Labeling about 75k episodes used Gemini 2.5 Pro and GPT-5.5. State checks have no independent ground truth, and the main text never measures what a weaker labeler would do to downstream success. Code, weights, and ARC-Trace are slated for release after review. A 1 Hz trace beside 15 Hz control means an action chunk can run on a judgment up to about a second old. The rise in language attention from 6% to 41% is, in the paper's own words, a diagnostic. The action expert reads those tokens. That does not by itself prove the behavior depends on them.

Terms

Source

What people are saying

Related papers

All paper explainers