AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems
Yiming Cheng, Alfin Wijaya Rahardja, Mengshi Zhang, Zihao Chen, Zhenpeng Chen, Yiling Lou
The University of Chicago / Fudan University / TensorBlock, Inc / Tsinghua University / University of Illinois Urbana-Champaign
cs.SE, cs.AI
2026-09-29
A multi-agent pipeline auto-reproduces real harness bugs from open-source agent repos at up to 28% success, building a 200-bug live benchmark where SOTA repair agents fix only 9%.
An agentic system is a backbone LLM wrapped in a harness: the software layer handling context management, tool integration, orchestration, and guardrails. Harnesses have grown large (OpenClaw ships about 200K lines), and bugs in this layer are now a meaningful source of agent failures. They are not ordinary software defects. They live in agent-specific structures, involve constant interaction with model providers, tools, and network services, and often surface only in particular execution states.
The measured numbers are rough: software repair agents fix harness bugs at 4.67% versus 40.67% on general software bugs. Progress is bottlenecked by data. The only executable benchmark, AgentIssue-Bench, holds 43 bugs and cost 150 manual hours, with a fixed size that invites training-data contamination. Automated reproduction pipelines such as SWE-Factory and SWE-bench-Live were built for general software and mostly break on agent repositories.
AgentBug-Smith is a fully automated multi-agent pipeline with three stages.
Reproduction success on 225 harness-bug issues, same pipeline across three backbone LLMs:
| Backbone | SWE-Factory | SWE-bench-Live | AgentBug-Smith |
| GPT-4.1-mini | 9.33% | 2.67% | 20.00% |
| Kimi-k2.5 | 8.44% | 2.67% | 20.44% |
| DeepSeek-v3.2 | 13.78% | 0.44% | 28.00% |
A 10.67 to 27.56 percentage-point lead at $0.56 to $2.44 per issue. Across all runs the pipeline reproduced 75 bugs, 45 of which neither baseline could reproduce, and manual inspection confirmed 74 of 75 (98.67%) as faithful reproductions. Identification holds up too: 95% accuracy on repositories and 92% on harness-bug issues, versus 35% and 46% for a vanilla LLM call.
The resulting Live-Harness-Bench holds 200 reproducible bugs: repositories average 123K lines of code, gold patches average 202.9 changed lines, 31.5% of bugs touch tool registries and action interfaces, 30.5% touch context and memory management, and the newest issues date to August 2026.
Three mainstream repair agents (all on GPT-4.1-mini) on the benchmark:
| Agent | Plausibly resolved | Correctly resolved |
| mini-SWE-agent | 19.50% | 9.00% |
| OpenHands | 17.50% | 8.50% |
| AutoCodeRover | 6.00% | 3.50% |
"Plausibly resolved" means the patch passes the reproducing tests; "correctly resolved" adds manual confirmation of semantic equivalence with the developer's patch. Against the 40.67% these agents report on SWE-Bench Verified, harness bugs are a different difficulty class. Used as a knowledge base (121 train, 79 test, no repository overlap), a distilled textual SKILL.md lifts mini-SWE-agent's correct resolution from 1.27% to 7.59%, with function-level localization rising from 34.60% to 66.41%.
This is the first executable harness-bug benchmark that grows with upstream development. The manual route cost 150 hours for 43 bugs; this pipeline pays a few dollars per issue, has already stacked up 200, and keeps ingesting new GitHub issues, pushing the contamination window forward.
The 9% versus 40.67% contrast settles a question: repair agents tuned on general software collapse on harness code. Claims that agents can fix agent infrastructure need this kind of benchmark, not a generic SWE score.
The skill-distillation experiment sketches a loop: real failures and fixes flow into the benchmark, and the benchmark feeds repair agents. The bigger it grows, the more the loop holds.
Reproduction tops out at 28%. Unsuccessful runs number 162 to 180 per backbone, dominated by environment failures: dependency installation, Docker configuration, container disk exhaustion. The authors acknowledge that failure distributions shift across backbone models, with no general fix. Evidence for joint optimization is a single AgentScope #1297 trace, which the paper itself calls a representative execution record, not a causal ablation.
The downstream experiments are small. The distillation test set is 79 bugs, the improvement starts from a near-floor 1.27%, and 7.59% is still low in absolute terms, verified on mini-SWE-agent alone. Also, the baselines were never designed for harness bugs, so beating them shows general techniques do not transfer; how far the pipeline sits from practical use is answered by the 28% rate itself.