Adaptive Consistency Graph lifts long-horizon agent success from 44.5% to 50.2%

Beihang · hf · 2026-09-29

Long-horizon LLM agents drift from original goals as task requirements, historical evidence, and execution state gradually disconnect across long sequences of dependent actions. Beihang's Adaptive Consistency Graph (ACG) incrementally organizes execution evidence and provenance in a persistent graph, then builds a temporary requirement-centered view per decision under a bounded context budget—without replacing the base agent's planner or tool executor.

In matched evaluation, ACG raises GPT-5.6-luna's average success from 44.5% with ReAct to 50.2%, with the largest gain on BrowseComp-Plus (73.5% vs 62.4%). The paper also analyzes trajectory structure and inference cost.

Original post →

More from coding & agent

coding & agent channel →