Agent reliability degrades with trajectory length, but no benchmark isolates it, dev finds

rio_ARC · reddit · 2026-09-04

A developer reports that 5–10 step agent workflows look stable while longer trajectories show distinct failure modes: unnecessary replanning, early mistakes propagating, decaying context, and retries that raise cost without improving outcomes. They want quantified curves of success rate, tool-call accuracy, recovery rate, cost per successful task and human intervention as a function of trajectory length. After surveying LangSmith/LangGraph, Lyzr Agent Studio, CrewAI and Letta, they found no benchmark isolating trajectory length, and ask the community for prior experiments.

Related event: Reddit Debates: When Do Long Agent Trajectories Break Down?(2 posts)→

Original post →

More from coding & agent

coding & agent channel →