AdaMAST: Automating Agent Failure Taxonomies Boosts SWE-bench to 70.7%
berkeley_ai · x · 2026-08-01
Researchers from UT Austin released a new paper on AdaMAST (Adaptive Failure Taxonomies), introducing a method to automatically categorize how and why AI agents fail.
- Core Mechanism: It learns system-specific failure taxonomies from raw execution traces and reuses them across runs to identify exactly where, when, and how agents fail.
- Performance: When integrated as a Claude Code skill, it improved the SWE-bench Verified Mini score from 64.0% to 70.7%.
- Benchmarks: On TerminalBench 2, it achieved 89.9% when using AdaMAST-generated taxonomies with an Opus 4.6/Forgecode harness and a Best-of-N Judge.
Related event: AdaMAST Automates Agent Failure Classification(2 posts)→
More from coding & agent
- Supermemory Launches Cross-Tool MCP to Share Long-Term Memory Among AI Agents — julianweisser · 2026-08-01
- Google Uses AI to Accelerate Vulnerability Patching, Fixing Hundreds of Chrome Security Bugs — steren · 2026-08-01
- Qwen Releases UI-Agent Technical Report: Unifying Cross-Platform GUI Interaction — _akhaliq · 2026-08-01
- Testing Xcode 27 Coding Agent: Handles Deployment but Fails Complex Game Logic — atShruti · 2026-08-01
- JAX Introduces Custom Types: Enabling Differentiable Rasterization Beyond Tensors — srush_nlp · 2026-08-01
- Hermes Desktop Launches Native Plugin SDK with Hot-Reloading and Deep UI Customization — NousResearch · 2026-08-01