New Benchmark Shows LLMs Struggle to Understand Agent Failures
jiank_uiuc · x · 2026-08-15
Amidst the hype of recursive self-improvement, researchers introduced Who&When Pro, a large-scale benchmark with over 12k failed trajectories spanning 26 benchmarks and 15 agent frameworks.
Key Findings:
- Localization: Current LLMs can sometimes locate where a failure happened.
- Attribution: Identifying the underlying type of failure remains significantly harder.
- Gap: This reveals a major gap for agents expected to autonomously diagnose and learn from their mistakes.
More from Research
- OpenMed Open Source Solution Solves Medical Data Privacy Challenges — aigclink · 2026-08-15
- Paper proposes efficient approximation for KL divergence between discrete normal distributions — FrnkNlsn · 2026-08-15
- RL Conference 2026 to focus on agents and self-improvement — tw_killian · 2026-08-15
- Fall Detection System Using Pose Estimation and Deep Learning — rsasaki0109 · 2026-08-15
- Study reveals hybrid LLMs organize computation around expensive full-attention layers — burny_tech · 2026-08-15
- New Framework Enables Multimodal Control for Humanoid Robots via Video and Music — zhengyiluo · 2026-08-15