New Benchmark Shows LLMs Struggle to Understand Agent Failures

jiank_uiuc · x · 2026-08-15

Amidst the hype of recursive self-improvement, researchers introduced Who&When Pro, a large-scale benchmark with over 12k failed trajectories spanning 26 benchmarks and 15 agent frameworks.

Key Findings:

Original post →

More from Research

Research channel →