3-layer MLP trained on 100 success traces beats GPT-5 at finding agent failure steps, 5000x faster

Tracing Agentic Failure from the Flow of Success

Samuel Yeh, Yiwen Zhu, Shaleen Deep, Sharon Li

cs.AI, cs.CL

2026-07-14

A one-class Neural CDE trained on 100 success traces flags agent failure steps by anomaly score: +20 F1 over GPT-5 in-domain, +7 OOD, 200-5000x faster, zero tokens.

What problem this solves

When an LLM agent system fails a long-horizon task, the trajectory can span dozens to hundreds of steps, and later actions often partially mask the earlier mistake. The Who&When benchmark put state-of-the-art reasoning models below 15% accuracy on localizing the error step, and manual tracing costs an expert hours per trajectory. Existing fixes split into two camps, both with hard limits: prompting pipelines burn frontier-LLM tokens and seconds per diagnosis, and RL-post-trained attributors need step-level error labels on failure trajectories, which are expensive and inherently ambiguous, since deciding which step doomed the run assumes a correction oracle nobody has in practice.

OAT changes the problem setting: train only on successful trajectories, which every deployed system produces for free, and localize error steps at inference time.

Method

The framing is one-class learning in representation space: learn what the normal flow looks like and flag deviations.

Results

In-domain (MCP-Atlas, 88 author-annotated failure trajectories):

MethodF1Hit rate
GPT-4o0.2120.250
GPT-50.1810.250
OAT (top-k)0.4200.777
OAT (CP)0.4350.566

Out-of-distribution (Who&When, 184 GPT-4o multi-agent trajectories): OAT top-k reaches 0.225 F1 and 0.451 hit rate against 0.152 and 0.275 for GPT-5, roughly +7 F1. AUROC is 0.629 in-domain and 0.758 OOD.

Efficiency is the bigger gap: OAT scores a trajectory in about 7 ms with zero output tokens and under 1 GB of VRAM; GPT-4o averages 4.2 s and GPT-5 about 39.6 s plus 3,012 output tokens, a 200–5000x range.

Ablations: swapping the CDE for a Neural ODE or an RNN hurts across the board; the gate buys +0.172 AUROC OOD at a 0.028 in-domain cost; extracting representations with proxy LLMs (Llama-4-Scout, Gemma-4, GPT-oss-120B) instead of the generator costs little, so the generator's internals are not required; later layers give stronger signal than earlier ones.

Why it matters

This is a failure-diagnosis route cheap enough to leave running. Prompting pipelines cost thousands of tokens and tens of seconds per diagnosis, which limits them to spot checks; OAT trains on 100 success trajectories and scores in milliseconds, so it can sit inside a production system as a real-time monitor, with per-step anomaly scores doubling as an observability signal. The honest framing: 0.435 in-domain F1 makes it a coarse filter that ranks which steps a human should read first, not a replacement for reading them.

Limitations

The authors' own case studies show the model misses early, latent errors buried in noisy reasoning text (statements like "no suitable tool" also appear in successful trajectories), and the hard conformal threshold drops causally significant steps that land just below the cut, which happened in one analyzed case.

The evaluation also deserves discounts. Scale is small: 103 success and 88 failure trajectories in-domain, with failure-step labels annotated by the authors themselves. The prompting baselines are weak in this setup: a random-step baseline scores 0.255 F1 in-domain, above both GPT-4o and GPT-5, so part of OAT's margin comes from the baselines being bad rather than from absolute strength. The OOD pipeline also includes a CORAL second-order alignment step disclosed only in the appendix; how much survives without it is not reported separately.

Terms

Source

What people are saying

Related papers

All paper explainers