Harness-R1: Agents Learn from Failure Trajectories to Patch Themselves

dair_ai · x · 2026-08-05

DAIR.AI introduced Harness-R1, a novel approach for self-improving agents that leverages historical failure trajectories. It post-trains a dedicated 9B 'engineer' model using online RL to convert batches of target-agent failures into validated executable runtime patches.\n\nBy only updating the engineer model, the target agent avoids drifting under the reward signal. In tests across WebShop, ALFWorld, and DBBench, this method lifted vanilla Qwen3.5-9B's success rate from 44.3% to 53.6%.

Original post →

More from coding & agent

coding & agent channel →