AgentBug-Smith paper: top coding agents fix just 9% of agent harness bugs
rohanpaul_ai · x · 2026-10-03
Researchers from UChicago, Fudan, Tsinghua and UIUC present AgentBug-Smith, an automated approach that discovers and reproduces real-world bugs in agent harnesses from open-source agentic systems, turning them into runnable tests and building a continuously growing Live-Harness-Bench (200 reproducible bugs so far).
Key findings:
- It beats general-software bug reproduction techniques by 10.67%–27.56% success rate across backbone LLMs.
- Harness bugs (tool calls, memory, prompts) depend on live model calls, making them hard to reproduce and test.
- The best of three coding agents fixed only 9% of these bugs, versus 40% on regular software bugs.
- A short guide distilled from past fixes lifted an agent from 1 to 6 correct fixes on 79 unseen bugs.
Practical takeaway: before trusting a coding agent with your agent's own code, test it on bugs you've already fixed.
More from coding & agent
- OpenClaw v2026.9.8 ships GPT-6.1 Sol support across 43 PRs, cuts memory use — heyneighbor · 2026-10-03
- Building personal productivity meta-tools is now fun: bash scripts have become full-blown apps — devenbhooshan · 2026-10-03
- K-Dense BYOK: Open-Source AI Co-Scientist Runs Locally With a Tamper-Proof Lab Notebook — RexDouglass · 2026-10-03
- Google paper: clean-context verifiers let explorer agents try wild proof ideas safely — solyarisoftware · 2026-10-03
- Building a Multiplayer Browser FPS Entirely With Claude: 8v8 Deathmatch, Free to Play — BlackberrySoft1082 · 2026-10-03
- VoidZero founders on TypeScript.fm: type-aware linting goes stable, Cloudflare deal and Vite's future — cnakazawa · 2026-10-03