UndoBench: Agents Hit 83.5% on Tasks but Only 46.7% at Fault Recovery

Dolly Sah · hf · 2026-10-06

A Hugging Face paper introduces UndoBench, a benchmark of 36 base workflows and 36 fault scenarios across 8 enterprise domains that decouples task competence from recovery capability via counterfactual paired trials with wire-level effect-history and state oracles.

Findings (5,760 executions / 2,880 paired trials across two open-weight models, two frameworks, three recovery paradigms):

Commercial API models reproduce the competence-recovery gap, showing that nominal-completion benchmarks mask critical recovery vulnerabilities.

Original post →

More from coding & agent

coding & agent channel →