AdaMAST: Boosting AI Agent Reliability via Failure Taxonomies

berkeley_ai · x · 2026-08-05

A UC Berkeley researcher introduced AdaMAST, a novel approach to improving AI agent reliability by teaching models how they fail.

The method learns system-specific failure taxonomies from raw traces and reuses them across runs to identify exactly where, when, and how agents fail.

Experiments show that using AdaMAST as a Claude Code skill boosts SWE-bench Verified Mini from 64.0% to 70.7%. On TerminalBench 2, it achieves 89.9% when combined with the Opus 4.6/Forgecode harness and a Best-of-N Judge.

Original post →

More from coding & agent

coding & agent channel →