Explicit belief states take first place on all four agent benchmarks, at 5x token cost

Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States

Yu Luo, Jiamin Jiang, Yimin Zuo, Xidao Wen, Rongchen Gao, Yongqian Sun, Shenglin Zhang, Guiyang Liu, Cheng Zhang, Fang Situ, Qi Zhou, Dan Pei

cs.AI

2026-10-01

PoS has agents maintain a validated belief state and escape zero-progress stalls; first on all four benchmarks and three backbones, up to +22.68% on ALFWorld, at 5x token cost.

What problem this solves

Long-horizon LLM agents mostly manage context as memory: keep the raw trajectory, compress it (ACON), tune granularity by relevance (PACE), or organize it under subgoals (HiAgent). All of these preserve evidence about what happened. None of them guarantees a coherent estimate of what the world looks like right now. As horizons grow, actions change the environment and new observations invalidate old inferences, so the context becomes a mix of stale facts and intermediate judgments. State updates written by the LLM itself can also contradict each other; the paper's example is a microwave recorded as both open and closed, which leaves the next action undefined.

A second failure mode gets a name here: Belief Trapping, where the agent keeps acting with no meaningful progress toward the goal, through repeated ineffective actions, loops, or collection of information it cannot use, until the budget runs out. Prior work either detects belief deviation during RL training and truncates trajectories (T3) or terminates unproductive gathering with an exhaustion gate (CGDP); neither restores progress inside the episode.

Method

PoS (Progression of States) is a pure inference-time framework: no training, no model swap, formalized on the POMDP belief state. Three parts:

The multiplicative form of H is deliberate. A geometric mean misses trapping when only one of stagnation or recurrence fires; a plain max flags any persistent gap as trapped even while the agent is making steady progress on it. The product fires only when a persistent unresolved gap coincides with stagnation or recurrence.

Results

Four benchmarks (ALFWorld and LOCA-Bench for execution; RCA-100 for microservice root-cause analysis; ClinDiag for clinical diagnosis), three backbones (Qwen3.7-Plus, Kimi-K3, GLM-5.3), against Raw Trajectory, ACON, PACE, HiAgent, and LongHorizon-Harness. PoS ranks first on all 12 benchmark-by-backbone combinations. Relative gains over the strongest same-backbone baseline reach 22.68% on ALFWorld, 7.53% on LOCA-Bench, 37.89% on RCA-100 joint accuracy, and 11.31% on ClinDiag.

Setting (Qwen3.7-Plus)Raw TrajectoryBest baselinePoS
ALFWorld success62.6972.39 (LongHorizon-Harness)88.81
LOCA-Bench success43.6252.57 (LongHorizon-Harness)56.38
RCA-100 joint accuracy24.2728.16 (PACE/HiAgent)38.83

Ablations: dropping consistency validation costs 14.93 points on ALFWorld and 11.81 on LOCA-Bench, but only 0.33 to 0.66 on ClinDiag; dropping trapping diagnosis and recovery costs 4.85 to 6.79 points on RCA-100 and 2.65 to 3.97 on ClinDiag. Replacing the factorized recovery with a generic recover prompt degrades all four benchmarks; ALFWorld falls from 88.81 to 76.87.

On cost and scale: per-episode tokens on RCA-100 rise from 355.7K to 1800.1K (5.06x), with the task agent itself consuming 20.9% less than raw; the overhead sits in belief construction (697K) and auditing (822K). As LOCA-Bench context grows from 96K to 256K, PoS stays roughly flat and leads the best baseline by 10.67 to 16.00 points at 256K. Trapping is common, hitting 14% to 94% of episodes depending on setting, and lower incidence does not imply higher accuracy: on RCA-100, Kimi-K3 traps in 74% of episodes versus 53% for GLM-5.3, yet Kimi-K3 scores higher joint accuracy.

Why it matters

The paper's read on the agent-memory line is blunt: for modern long-context models, compressing or reorganizing history does not consistently beat keeping it raw. PACE sits 24.57 to 28.19 points below raw on LOCA-Bench, and on ClinDiag with GLM-5.3 no context-management baseline beats the raw trajectory at all. The compute belongs in maintaining a checkable estimate of the current world, a shift from storing history to tending state. PoS spends 5x tokens to buy accuracy and stability at long contexts, and reports the tradeoff explicitly enough for an engineering decision. The mechanism needs no training and no stronger auxiliary model; the same backbone audits itself, so teams doing root-cause analysis, diagnostic dialogue, or embodied tasks can adopt the pattern directly.

Limitations

Terms

Source

Related papers

All paper explainers