AdaMAST: Boosting AI Agent Reliability via Failure Taxonomies
berkeley_ai · x · 2026-08-05
A UC Berkeley researcher introduced AdaMAST, a novel approach to improving AI agent reliability by teaching models how they fail.
The method learns system-specific failure taxonomies from raw traces and reuses them across runs to identify exactly where, when, and how agents fail.
Experiments show that using AdaMAST as a Claude Code skill boosts SWE-bench Verified Mini from 64.0% to 70.7%. On TerminalBench 2, it achieves 89.9% when combined with the Opus 4.6/Forgecode harness and a Best-of-N Judge.
More from coding & agent
- OpenAI and Anthropic AI Agents Attacked Real Systems in Cyber Tests — jedisct1 · 2026-08-05
- Minimalist AI Coding Harness Boosts Performance and Cuts Costs — zainhas · 2026-08-05
- PosterMELD: Multi-Agent System for Paper-to-Poster Generation — Haojie Hu · 2026-08-05
- Migrating from Single Provider API to Aggregation Gateway: Developer Shares Lessons Learned — Loud_Ice4487 · 2026-08-05
- Sakana AI's Agentic Systems Enter Production at Daiwa Securities — tkasasagi · 2026-08-05
- Testing MiniMax H3 and Others for Music Video Creation with Open-Source Tool Velorn — VisualFXMan · 2026-08-05