Harness-of-Harness lets coding agents run unsupervised for days, 52% avg gains
dair_ai · x · 2026-09-02
New research on Harness-of-Harness wraps whatever coding harness you already run and organizes executions into repeated planning-coding-testing increments, enabling agents to keep building for days without a human.
Key mechanics:
- Balances repair against capability growth
- Scopes work into small verifiable steps
- Keeps implementation-time testing separate from independent evaluation
- Constrains outputs rather than the workflow
Across GameCraft-Bench, FrontierSWE and ProgramBench with three harness-model pairs, it averages a 52.25% relative gain over standalone harnesses after three iterations, peaking at 82.86%.
Related event: Harness-of-Harness Lets Coding Agents Run Unattended for Days(2 posts)→
More from coding & agent
- Developer asks: is LangChain still worth it vs rolling your own agent harness? — curious_vii · 2026-09-03
- Inference Engineering Is Just a Recipe: vLLM/SGLang, Replicas, Cache-Aware Routing — GabGarrett · 2026-09-03
- Developer accidentally built an entire agent factory with Fable 5.1 — 0xkarasy · 2026-09-03
- Microsoft adds Fabric data agents to Foundry agents via Fabric IQ (preview) — adnan_hashmi · 2026-09-03
- Databricks pitches agent-native data infrastructure, Lakebase Postgres at VLDB 2026 — matei_zaharia · 2026-09-03
- doodlestein ships a comprehensive web app review skill after months of debugging — doodlestein · 2026-09-03