AdaMAST: Boosting AI Agent Performance with Failure Taxonomies
abeirami · x · 2026-07-31
Developer @mertcemri introduced AdaMAST, a method that learns system-specific failure taxonomies from raw execution traces. It reuses these insights across subsequent runs to identify exactly where, when, and how AI agents fail.
In testing, integrating it as a Claude Code skill boosted the SWE-bench Verified Mini pass rate from 64.0% to 70.7%. On TerminalBench 2, using AdaMAST-generated taxonomies with an Opus 4.6/Forgecode harness and a Best-of-N Judge achieved a high score of 89.9%.
More from coding & agent
- Marble (YC S26) Launches: Uses CV for Inventory and AI Agents for Restaurant Ops — ycombinator · 2026-07-31
- Developer Runs Claude for 24 Hours Straight to Build Open-Sourced 3D Town — Dr_Singularity · 2026-07-31
- MCP Reliability Sidecar Slashes Harmful Events Across 200 Agent Runs — FewScarcity6957 · 2026-07-31
- Solo Dev Vibe-Codes Full 3D Game End-to-End Using Grok — Daniel_Farinax · 2026-07-31
- CodeGraph 1.0 Fixes Symlink Traversal Bug Preventing Agent Secret Leaks — JeremyCMorgan · 2026-07-31
- Gradio Teases Upcoming Resume Sessions for Long-Running AI Jobs — Gradio · 2026-07-31