UndoBench: Agents Hit 83.5% on Tasks but Only 46.7% at Fault Recovery
Dolly Sah · hf · 2026-10-06
A Hugging Face paper introduces UndoBench, a benchmark of 36 base workflows and 36 fault scenarios across 8 enterprise domains that decouples task competence from recovery capability via counterfactual paired trials with wire-level effect-history and state oracles.
Findings (5,760 executions / 2,880 paired trials across two open-weight models, two frameworks, three recovery paradigms):
- Nominal competence reaches 83.54% while conditional recovery success rate falls to 46.72%.
- Naive retry duplicates external effects in 53.33% of trials.
- Recovery is phase-dependent: methods tie before mutation; naive retry, per-call idempotency, and zero-privilege journaling collapse during partial mutation; verification and server-side idempotency help most post-commit pre-acknowledgment.
Commercial API models reproduce the competence-recovery gap, showing that nominal-completion benchmarks mask critical recovery vulnerabilities.
More from coding & agent
- Kapa MCP lets you set up a docs agent, Slack bot and PR fixes from one chat in Claude or Cursor — CShorten30 · 2026-10-06
- Heavy Claude Code User Seeks Open-Source Agent Harness with Model Routing — lulz_lurker · 2026-10-06
- Harness Engineering paper breaks down how Claude Code, Codex and Gemini CLI are built — joemeno · 2026-10-06
- Jev-as-a-Judge: using the new Jev model to boost agent eval reliability — omarsar0 · 2026-10-06
- Maven launches free AI Builders crash course: from building LLMs to agents and production — leslysandra · 2026-10-06
- Experiment shows LLM session history can override skill-file rules even after correction — KhuyenTran16 · 2026-10-06