Agent reliability degrades with trajectory length, but no benchmark isolates it, dev finds
rio_ARC · reddit · 2026-09-04
A developer reports that 5–10 step agent workflows look stable while longer trajectories show distinct failure modes: unnecessary replanning, early mistakes propagating, decaying context, and retries that raise cost without improving outcomes. They want quantified curves of success rate, tool-call accuracy, recovery rate, cost per successful task and human intervention as a function of trajectory length. After surveying LangSmith/LangGraph, Lyzr Agent Studio, CrewAI and Letta, they found no benchmark isolating trajectory length, and ask the community for prior experiments.
Related event: Reddit Debates: When Do Long Agent Trajectories Break Down?(2 posts)→
More from coding & agent
- POSTECH's PACE uses coordinated agents to surface hidden conflicts in user requests — POSTECH · 2026-09-04
- Astra solves Excel World Championship cases ~4x faster than human champions using pure computer use — sandersted · 2026-09-04
- Dev builds a Magicka-inspired MMO in just 3 days using fable 5.1 — TAbrodi · 2026-09-04
- Codex Lacks /insights, So a Dev Built an Open-Source Local Analytics Plugin — Salt_Hyena5896 · 2026-09-04
- Creator assigns overnight tasks to AI agents, hopes for renders not a coup — bilawalsidhu · 2026-09-04
- Amazon-Microsoft paper: skill-guided action chunking lifts agent success to 67.2% while halving LLM calls — rohanpaul_ai · 2026-09-04