Harness evolution hits a ceiling: LoRA cuts agent content failures from 1/4 to 1/20
rohanpaul_ai · x · 2026-10-11
The Stanford paper (arXiv:2610.11655) systematically answers whether to evolve an agent's harness or train its weights: the right lever can be read off the failure composition.
- Label failed trajectories by the first firing signal to separate process failures (blocked calls, loops, exhausted step budgets) from content failures (poor delivered plans). Harness evolution repairs the former; content failures are what weight training is for.
- On DeepPlanning, a self-evolving harness loop lifted Qwen3.5-4B's held-out score 0.16→0.30 and 9B 0.32→0.44; 4B delivery rose 55%→90%.
- LoRA adapters trained on evolved-harness trajectories internalize gains: +0.13 under the original harness for both sizes; on 4B they stack with the harness to more than double held-out scores; on 9B the adapter alone matches the full evolution line, cutting content failures from a quarter to one in twenty. A placebo adapter on answer-shuffled data falls below base, confirming the effect. The loop also transfers to WebArena-Lit.
Related event: Stanford study finds harness evolution hits a ceiling for agents(2 posts)→
More from coding & agent
- Instead of hoarding 43 browser tabs, let agents build a TLDR digest and save them for later — RachelVT42 · 2026-10-11
- Researcher builds keyboard-driven macOS PDF reader with Claude, open-sources it on GitHub — sethlazar · 2026-10-11
- Grok Bot agent gets X-wide job assignments: 10 copy-paste routines from daily launch radars to free-credit alerts — dr_cintas · 2026-10-11
- 16 parallel Cursor projects later, the most-used prompt is just "catch me up on everything" — RachelVT42 · 2026-10-11
- Multi-agent in production: what works, what gets abandoned, and who controls the flow — fkj83 · 2026-10-11
- Open-source MCP server lets Claude Opus 5.5 write complete FL Studio tracks end to end — delarein · 2026-10-11