CMU paper: RL-trained 4B proposer edits agent harness code, beats its 35B teacher
omarsar0 · x · 2026-10-04
A new CMU paper introduces "harness learning": improving agents by editing their harness code instead of their weights.
- An RL-trained proposer model reads a task, the current harness, and an execution report, then writes a code edit to the harness
- The reward is the revised harness's score; the solver model never changes
- The trained 4B proposer beats its 35B teacher at single-step revision on Reasoning Gym, including task families unseen in training
- A proposer trained on HotpotQA keeps improving harnesses on MuSiQue and 2WikiMultihopQA
The authors highlight editing harness code as an emerging AI engineering skill, a shift they see accelerating.
More from coding & agent
- Vercel hits $600M annualized revenue, up 148%, as coding agents drive half of new business — evilrabbit_ · 2026-10-04
- Hallmark: an open-source skill making Claude Code, Cursor and Codex UIs look less AI-generated — tom_doerr · 2026-10-04
- Stop fine-tuning to fix retrieval problems: Oracle technologist on where knowledge should live — AI Engineer · 2026-10-04
- KMP: recovering project decisions and evidence across Claude and Codex via MCP — Mountain-Raise-4556 · 2026-10-04
- (Lean)DOOM: DOOM fully rewritten in the Lean proof assistant, with formal proofs included — akbirthko · 2026-10-04
- Jin: a minimalist coding agent that swaps MCP/plugins for prompts and bash — aldapsiger · 2026-10-04