LoongReflect: Enhancing Long-Horizon Reflection in Search Agents via Global Distillation

_reachsumit · x · 2026-08-13

To address the challenge of local-global mismatch in LLM agents' long-horizon reasoning, the paper introduces LoongReflect, a novel training framework.

It formulates reflection as a memory-control policy, allowing agents to perform explicit reflect and backtrack actions over a reversible trajectory tree. By combining look-ahead and extragradient-style coordination, the framework distills global perspectives into local reflective decisions, significantly improving the agent's performance in complex search tasks.

Original post →

More from coding & agent

coding & agent channel →