LongRCA Bench: Diagnosing Failures in Long-Horizon Agent Trajectories
Yunfei Zhang · hf · 2026-08-26
LongRCA Bench evaluates failure diagnosis across lengthy agent trajectories. The training-free RCTA method improves the attribution of responsible roles and root-cause steps.
More from coding & agent
- One audit prompt cut 170M tokens/month: cleaning up Hermes agent cron jobs — toddhooper · 2026-08-26
- Brian Armstrong shares a 3-step workflow for AI to mimic your writing style — ivan_bezdomny · 2026-08-26
- 8 Biggest Unsolved Problems in Evaluating AI Agents Today: Trajectories, Regressions, and Production Gaps — Alternative_Duck_908 · 2026-08-26
- Storing and tracking MCP inputs for reinforcement learning — frothyyyyyy · 2026-08-26
- Open-Source Guaardvark Simplifies ComfyUI with Voice Chat and MCP Integration — llama-of-death · 2026-08-26
- OpenAI: KV Cache is the largest and fastest-growing data structure in agentic inference — BenBajarin · 2026-08-26