AdaMAST: Boosting AI Agent Performance with Failure Taxonomies

abeirami · x · 2026-07-31

Developer @mertcemri introduced AdaMAST, a method that learns system-specific failure taxonomies from raw execution traces. It reuses these insights across subsequent runs to identify exactly where, when, and how AI agents fail.

In testing, integrating it as a Claude Code skill boosted the SWE-bench Verified Mini pass rate from 64.0% to 70.7%. On TerminalBench 2, using AdaMAST-generated taxonomies with an Opus 4.6/Forgecode harness and a Best-of-N Judge achieved a high score of 89.9%.

Original post →

More from coding & agent

coding & agent channel →