Long-running agents: error compounding, context pollution and weak self-correction
Sad_Lavishness_53 · reddit · 2026-10-05
A structured Reddit discussion on why agents collapse on long tasks even when every single step is easy:
- Error compounding: at 95% per-step accuracy, a 20-step chain succeeds only 36% of the time (0.95^20)
- Context pollution: after enough tool calls, the context fills with stale outputs, dead ends and failed attempts, and the model drifts from the original goal
- Weak self-correction: when a step goes wrong, models tend to build on the error instead of backing up
The core open question: is the fix better models, or better scaffolding — checkpoints, verifier steps, state stored outside the context window? The author asks for practitioner input on history pruning vs. full context, whether separate critic/verifier models actually pay off, and where agents typically start breaking in real tasks.
More from coding & agent
- W&B shows how to turn a production agent failure trace into an eval — wandb · 2026-10-06
- Dev builds Claude Code skill that writes better HTML plans with plain language, mockups and linting — trq212 · 2026-10-06
- OpenAI ships compaction in Responses API, sparking vendor lock-in debate among developers — pvncher · 2026-10-06
- DoorDash launches MCP and CLI for agentic ordering; dev auto-restocks office pantry with camera — Scobleizer · 2026-10-06
- Anthropic ships Claude Code mods: TypeScript functions that rewrite prompts and UI — thione · 2026-10-06
- OpenAI launches Codex Security Cloud for scheduled full-repo GitHub security scans — thione · 2026-10-06