CAFE Improves Search Agents via Co-Evolving Agent and Critic
_reachsumit · x · 2026-08-26
Addressing the limitation of outcome-supervised search agents in localizing intermediate errors, the paper proposes CAFE (Coupled Agent--Feedback Evolution). The framework uses a shared-parameter model alternating between search-agent and critic roles, enabling the agent to request and use corrective feedback mid-trajectory. Combining online RL and offline preference optimization, CAFE allows the agent to learn feedback-guided recovery from its own failures. It outperforms RL-based search agents on average across seven benchmarks and retains gains on six out-of-domain benchmarks.
Related event: Tencent's CAFE Couples Search Agent and Critic for Self-Improvement(2 posts)→
More from Research
- How LLMs Self-Correct Mid-Generation: The Role of Reasoning RL and Instructions — dejanseo · 2026-08-26
- U. de Chile Students Publish Book on Maturana and Varela's Relevance in AI — PolarBearby · 2026-08-26
- Face Anything: 4D Face Reconstruction from Any Image Sequence (ECCV 2026) — rsasaki0109 · 2026-08-26
- Gemini 3.7 Flash helps revive interactive CMA-ES explainer site — doodlestein · 2026-08-26
- From PDE Numerical Solvers to Neural Emulators and Back: PhD Thesis — chaumian · 2026-08-26
- Combining PSGD-Kron and KL-Shampoo yields an optimizer without eig/inverse/solve — YouJiacheng · 2026-08-26