Overcoming State Loss in Long-Horizon Agents: New Framework Boosts Accuracy

Ziyu Ma · hf · 2026-08-04

Existing LLM agents often struggle with state tracking and error propagation due to ever-growing contexts during long-horizon tasks. Researchers introduce LongHorizon-Harness, reformulating execution as a task-state management problem.

At its core is the Manage-Execute-Audit (MEA) loop:

Experiments show significant accuracy improvements across benchmarks like WeaveBench, Terminal-Bench, and OSWorld. For instance, it raises Qwen3.7-Plus from 51.8% to 80.7% on WeaveBench and lifts Claude Opus 4.7 from 20.0% to 34.3% on an OSWorld subset.

Original post →

More from coding & agent

coding & agent channel →