Adaptive Failure Taxonomy Improves Coding Agents
AlexGDimakis · x · 2026-07-10
This work, which won Best Paper at an ICML Workshop, discusses the utility of an adaptive failure taxonomy:
- Can be used as a test-time scaling tool for best-of-N judges
- Can serve as a mutation feedback mechanism in optimization loops
- Can act as runtime feedback for coding agents
The author notes that past approaches relied on manually fixed failure taxonomies. In contrast, the new method observes agent rollouts to dynamically build a taxonomy tailored to the task and the agent's weaknesses. The paper reports that on Terminal Bench 2, combining Opus 4.6 / Forgecode harness with an adaptive taxonomy for best-of-N judging achieved 89.9%, a 15% improvement over fixed taxonomies.
Related event: Adaptive Failure Taxonomies Win ICML Workshop Best Paper(3 posts)→
More from coding & agent
- The browser main thread is expensive: a practical guide to JavaScript and CSS animation cost — jh3yy · 2026-09-11
- Inspired by OpenAI's 10,000-agent run, dev open-sources a crowdsourced agent problem-solving platform — Benjaminsen · 2026-09-11
- Lucid: open-source Mac app keeps your laptop awake only while AI agents run — Pitiful_Hedgehog_600 · 2026-09-11
- banteg's snail project crowdsources AI agents to finish matching Snail Mail's 20 remaining functions — banteg · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11