Harness-of-Harness paper: 52.25% avg gain across 3 benchmarks, builds playable FPS unattended
aigclink · x · 2026-09-05
The Harness-of-Harness paper (Shanghai AI Lab) details a meta-framework organizing coding-agent harnesses into planning-coding-testing loops with small verifiable increments and separated evaluation.
- Evaluated on GameCraft-Bench, FrontierSWE, and ProgramBench with Codex+GPT-5.5(high), OpenCode+DeepSeek-V4-Pro, and Pi+MiniMax-M3
- Average relative gain of 52.25% (max 82.86%) after three iterations
- Multi-day, 70+ iteration autonomous build of Fusepoint, a human-playable FPS with coherent story, mechanics, visuals, and audio, from a single PRD
Related event: Shanghai AI Lab's Harness-of-Harness builds a playable FPS autonomously(2 posts)→
More from coding & agent
- Building AI Agents: Memory, Tool Limits, and Orchestration Frameworks — mdancho84 · 2026-09-05
- A 7-step cheat sheet for building AI agents, from system prompt to evals — mdancho84 · 2026-09-05
- Claude Code v2.1.251 now saves effort levels per model via /effort or settings. — EricBuess · 2026-09-05
- Subagents Log Where Docs Fail: Docker Stuck Points and Install Gaps — lucasmeijer · 2026-09-05
- Dev Uses Subagents as Fake Users to Test His Coding Agent's Install Flow — lucasmeijer · 2026-09-05
- ChatGPT Android app hides the Codex menu but works via manually paired remote sessions — ezyang · 2026-09-05