arXiv: agent success follows a geometric decay law, collapsing within 16 steps
rohanpaul_ai · x · 2026-09-08
Paper [arXiv:2609.01660] measures long-horizon degradation in LLM agents across 9 models (open models 1.2B-671B plus 3 proprietary), 4 task families, 5 horizons, and 3 context regimes, analyzing 10,664 trajectories.
- Task success follows a geometric law governed by a single per-step reliability parameter; it rises with model scale but saturates well below 1, guaranteeing collapse at long horizons.
- On the agentic tool-use task, every model — including widely deployed systems — falls from near-perfect to near zero within 16 steps.
- Decay is driven by step count, not context length: bounding the context window steepened decay (logit slope -0.69 vs -0.44, p=3e-6), contradicting lost-in-the-middle explanations.
- Implication: benchmark pass rates don't prove production readiness; test at real workflow lengths and add checkpoints.
More from coding & agent
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- 105 hidden bugs, 2 repos: DeepSeek V4.1 Flash fixes 24 at $1.80 vs Opus 5's 27 at $51.33 — ChartsJournalX · 2026-09-11
- Investment Analyst Asks How to Build a Claude-Based Diligence Agent Stack — Careless_Tie2286 · 2026-09-11
- Treating agents like 50 First Dates: a 3-layer context system so every conversation doesn't start from zero — evielync · 2026-09-11
- Running the Firefox MCP on Android via Termux, ngrok, and mcp-proxy — Nervous-Strain7544 · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11