CAFE Improves Search Agents via Co-Evolving Agent and Critic

_reachsumit · x · 2026-08-26

Addressing the limitation of outcome-supervised search agents in localizing intermediate errors, the paper proposes CAFE (Coupled Agent--Feedback Evolution). The framework uses a shared-parameter model alternating between search-agent and critic roles, enabling the agent to request and use corrective feedback mid-trajectory. Combining online RL and offline preference optimization, CAFE allows the agent to learn feedback-guided recovery from its own failures. It outperforms RL-based search agents on average across seven benchmarks and retains gains on six out-of-domain benchmarks.

Related event: Tencent's CAFE Couples Search Agent and Critic for Self-Improvement(2 posts)→

Original post →

More from Research

Research channel →