A survey maps the real bottleneck in long-horizon agents

青稞AI · wechat · 2026-07-21

This long-form article summarizes a 149-page survey on long-horizon agents and argues that the key bottleneck for production agents is not just single-step intelligence, but sustained execution over long dependency chains.

Main idea

Long-horizon ability is defined by how well an agent can keep goals intact, manage context, recover from errors, and finish tasks over long stretches of interaction. The survey frames this as a joint evolution of:

Key framework

The article explains a three-level ladder:

It also traces the system shift from prompt engineering to context engineering and now to runtime harnesses that control tools, memory, workflows, validation, and recovery.

Why it matters

The piece argues that long-horizon agents fail mainly through goal drift, context corruption, sparse delayed rewards, and irreversible actions. In practice, this means agent competitiveness is increasingly a systems-engineering problem, not just a model-quality problem.

Original post →

More from coding & agent

coding & agent channel →