Adaptive Failure Taxonomy Improves Coding Agents
AlexGDimakis · x · 2026-07-10
This work, which won Best Paper at an ICML Workshop, discusses the utility of an adaptive failure taxonomy:
- Can be used as a test-time scaling tool for best-of-N judges
- Can serve as a mutation feedback mechanism in optimization loops
- Can act as runtime feedback for coding agents
The author notes that past approaches relied on manually fixed failure taxonomies. In contrast, the new method observes agent rollouts to dynamically build a taxonomy tailored to the task and the agent's weaknesses. The paper reports that on Terminal Bench 2, combining Opus 4.6 / Forgecode harness with an adaptive taxonomy for best-of-N judging achieved 89.9%, a 15% improvement over fixed taxonomies.
Related event: Adaptive Failure Taxonomies Win ICML Workshop Best Paper(3 posts)→
More from coding & agent
- A roundup of AI agents and MCP resources, including how to evaluate agents — _jaydeepkarale · 2026-07-21
- A full course shows how to build and deploy an AI agent with OpenAI and LangChain — _jaydeepkarale · 2026-07-21
- A beginner guide to AI agents points readers to a Stanford webinar — _jaydeepkarale · 2026-07-21
- A practical guide on how to evaluate AI agents — _jaydeepkarale · 2026-07-21
- MCP is headed toward easier scale, event-driven extensions, and workable file uploads — EricBuess · 2026-07-21
- Developers debate the missing composition model for AI agents — threepointone · 2026-07-21