No VLM, No VLA at Test Time: Code-Only Robot Policy Hits 70.24% on RoboDojo, 38.9 Points Over SOTA

Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement

Kairui Hu, Siyuan Hu, Fangzhou Hong, Zhaoxi Chen, Ziwei Liu

cs.RO, cs.CV

2026-10-09

COAP writes robot policies entirely in code that measures its own state, with no VLM or VLA at test time, reaching 70.24% on RoboDojo's 42 bimanual tasks, 38.9 points above SOTA.

What problem this solves

Most robot manipulation policies keep a model in the control loop. VLAs map camera images straight to actions, with the decision logic buried in their weights; Agent Harnesses such as Agent-as-Policy query a VLM at every decision and keep some state as text, but the action still comes out of a black box. In both cases the state is hard to inspect and hard to edit, every model call carries randomness, and improvement means new data and a training run that can wash out old tasks: sequential fine-tuning of π0.5 on five real-world tasks dropped its average score from 86.9 to 31.4. This paper flips the view: a robot plus its environment is an Embodied Turing Machine, where the world state is the tape and the policy is the rules. If code can measure that state accurately, every decision can be written as code, and no model is needed at runtime. COAP (Code-Only-as-Policy) is that idea built out.

Method

Explicit state. The program maintains one state in three parts: the environment part records each object's pose, size, class, and spatial relations, with error and source attached to every record; the robot part records joint angles, end-effector poses, and gripper openings; the task part records the current stage, retry counts, and whatever the task must remember, such as which color sits under each cup. The code is frozen at runtime, and a single program runs every episode of a task, including the held-out test layouts.

Perception is ordinary code. The perception layer is NumPy and OpenCV: project the robot mesh into the image from measured joint angles and mask it out, take connected regions that differ from the table color, intersect view rays with the table plane to localize objects, then fit each object's known 3D shape over position and yaw and keep the best-overlap match. Camera parameters, robot meshes, and object sizes are constants published with the benchmark, part of the rules; which objects are present and where is state, measured fresh from images each episode. Occluded objects keep their last recorded pose, and millimeter-level tasks re-measure with a wrist camera before the final approach.

A shared library, improved offline. The library has four layers: task programs, object families, manipulation primitives, perception and motion. A new task composes existing modules, inherits an old task and overrides one decision, or adds a default-off parameter to a shared function, with a trace-equivalence condition guaranteeing old tasks keep their behavior. Across the 42 tasks, 83% of the code a task runs comes from the library. Development is a closed loop run by Claude Opus 5.5 coding agents, no human in the loop and no access to the test set: an agent reads the recorded state of failed episodes, pinpoints the variable at fault, and proposes a diff; branches pass CPU checks, then run in parallel on validation layouts, and a diff merges only if the success count clears a margin above rerun noise. Every branch is a complete runnable program that needs no training, so variants of one decision (grasp a cup by rim, handle, or top) run side by side, like a breadth-first search over code logic.

Results

RoboDojo is a 42-task bimanual benchmark covering memory, precision, open instruction following, generalization, and long-horizon manipulation, three seeds of 50 layouts per task. Inputs are three RGB streams, proprioception, and the instruction; no depth, no segmentation, no model service at test time.

MethodTypeOverall SR
PhysicalRSI (prior best)Agent Harness31.38%
Awomo-0.5 (best WAM)World Action Model29.64%
Simate-beta (best VLA)VLA27.96%
COAP (this paper)code only, no model at test70.24%

COAP sits 38.9 points above the state of the art and leads all five dimensions. The gaps are widest on Memory (89.9%, +39.6 over the best baseline) and Precision (75.1%, +41.6), the two dimensions that test stored values and measurement accuracy. Generalization is 65.7% (77.3% on standard layouts, 54.1% on randomized ones); the smallest margin is Long-Horizon at 56.4%, +13.0.

The mechanism-level numbers are more telling. The same program succeeds in 63% of episodes on measured state and 83% on simulator ground truth, so the ceiling sits at state measurement. A control step takes a median 0.3 ms, and an episode uses 0.8–2.4% of a VLA's compute and under 0.1% of an Agent Harness's. In episodes where something went wrong, the program noticed the problem itself 76% of the time and re-attempted with a new method 34%. New tasks built from scratch reach 47% average success within 15 hours with no data and no training; the strongest VLA reaches 4.7% on the same tasks after training on demonstrations.

Why it matters

For anyone building robot systems, this is a direct reminder: on tasks where state can be measured and structured, try code before reaching for a model. Explicit state plus explicit logic wins Memory and Precision by 40 points at a few percent of the compute. Three practical slots:

The boundary is equally clear: the whole approach assumes measurable state, and in open worlds that assumption often fails.

Limitations

The authors' own list: everything is in simulation, and perception leans on the benchmark's published camera parameters, robot meshes, and object sizes, all of which need measuring on a real robot, with contact parameters retuned; offline development costs time for every new task; generalization is limited, with randomized layouts 23 points below standard ones, and hard-coded values appearing as real failures (a fixed size threshold filtered out a small watch).

Three things also look shaky on close reading. The perception pipeline needs each object's known 3D shape and a known table color, so this is a known-objects, known-table setting; deformable objects are an explicit weak spot, with the paper's own example being a shirt hem slipping from the gripper. 62% of failures trace to state measurement built on classical geometric vision, and the paper's answer is only to fix the perception code, with no evidence of how far that carries. The comparison is also asymmetric: baselines are general models, while COAP is a per-task program developed and debugged offline by coding agents plus a shared library, so it holds as a paradigm argument but not as a blanket claim that code beats VLAs; the compute cost of the development loop is not reported.

Terms

Source

Related papers

All paper explainers