A 149-page survey says long-horizon agents depend on harnesses, not just bigger models

机器之心 · wechat · 2026-07-25

Long-horizon agents need more than long context

This 149-page survey argues that long-horizon agent capability is a system-level property, not just a stronger base model. It reviewed 900+ papers and engineering systems, and frames progress as the co-evolution of external harness engineering and internal model optimization.

What makes long tasks hard

The paper’s core framework

It separates the field into three task levels and three capability levels:

Three phases of agent control

Harness components

The survey breaks the harness into six pieces:

Internal optimization directions

It also surveys training-side work across:

Why it matters

The paper’s main takeaway is that long-horizon agents will be judged less by how long they can run, and more by whether they can remain effective, economical, and trustworthy over long dependency chains. It highlights open problems in continuous learning, real-environment evaluation, budget-aware execution, and governance-by-harness.

Original post →

More from AGI Musings

AGI Musings channel →