Microsoft & Tsinghua: structured run views lift GPT-5.1 agent failure localization from 3.6% to 31.4%

rohanpaul_ai · x · 2026-09-09

A Microsoft–Tsinghua paper shows that diagnosing LLM agent failures improves dramatically when the judge model gets a structured view of the run instead of the raw conversation.

Key findings

The authors note the problem remains unsolved at 31%, and recommend building structured traces and explicit failure checks into systems before asking an LLM to explain failures.

Related event: Microsoft-Tsinghua Paper Boosts Agent Failure Diagnosis with Structured Run Views(2 posts)→

Original post →

More from coding & agent

coding & agent channel →