AgentBug-Smith paper: top coding agents fix just 9% of agent harness bugs
rohanpaul_ai · x · 2026-10-03
The author shares the arXiv link for AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems, by researchers from UChicago, Fudan, TensorBlock, Tsinghua and UIUC.
The method automatically reproduces real-world agent harness bugs, outperforming existing general-software bug reproduction techniques by 10.67%–27.56% across backbone LLMs. It yields Live-Harness-Bench, a live benchmark of 200 reproducible harness bugs. Systematic evaluation shows state-of-the-art software agents struggle to repair real harness bugs, while distilling lessons from past fixes significantly improves repair performance.
More from coding & agent
- Fable 5.5 hallucinates less and replicates small models like Jev in hours, claims Bindu Reddy — bindureddy · 2026-10-03
- Anthropic engineer: future models will get much better at code deletion and simplification — simpsoka · 2026-10-03
- The Best AI Workflows Keep Friction Exactly Where Mistakes Matter — alifcoder · 2026-10-03
- Dev builds browser 3D game from scratch with Opus 5.5, Blender and Three.js — jason_mayes · 2026-10-03
- Cloudflare Durable Objects now survive client disconnects for long-running agents — threepointone · 2026-10-03
- Dev tired of nbviewer crashing builds serverless browser Jupyter notebook renderer, MIT-licensed — cneuralnetwork · 2026-10-03