Harness-R1: Teaching Agents to Self-Repair via Failure Trajectories
alex_verem · x · 2026-08-05
Core Contribution of Harness-R1
The paper introduces Harness-R1, a novel method that uses reinforcement learning to enable agents to learn from their failed interaction trajectories. It automatically edits the executable runtime harness to facilitate self-repair and performance enhancement.
Mechanism
- Role Separation: Introduces a dedicated 9B-parameter "engineer model" that analyzes the target agent's failures and generates validated executable patches.
- Online RL: The engineer model is trained using outcome-based rewards (whether the task succeeds upon rerun) to optimize its editing policy.
- Performance: Across benchmarks like WebShop and ALFWorld, this approach increased the task success rate of a vanilla Qwen3.5-9B from 44.3% to 53.6% (+9.3 percentage points).
Related event: Harness-R1 Enables Agents to Self-Repair from Failure Trajectories(3 posts)→
More from coding & agent
- Wii 3D game ported to Nintendo DS via ChatGPT, using just 9% of weekly quota — amplifiedamp · 2026-09-22
- AI memory system built with Jev claims 10x speed, 6x lower cost — realferrari · 2026-09-22
- Claude usage is shifting from 'write this' to context-heavy workflows — Thick-Session7153 · 2026-09-22
- riderless: A Zero-Token Decision API on Gemma 4 Hits 93.1% Accuracy on One RTX 5090 — Pale-Soil-2524 · 2026-09-22
- Cursor's Composer 1 powers a new AI-native coding interview: build a real project in 1 hour — steipete · 2026-09-22
- Theo laments $200/month AI coding subs no longer delivering 'practically unlimited' usage — haydendevs · 2026-09-22