AnySearch: RL Framework Adapts Single Search Policy to Any Budget

_reachsumit · x · 2026-09-02

The paper 'One Policy, Any Budget' proposes AnySearch, a framework enabling a single LLM search agent policy to adapt to any budget constraint via curriculum reinforcement learning. Existing methods train under fixed budgets and fail to adapt when constraints vary. The approach has two phases: first, training with explicit budget state injection and structured prompts; second, removing the scaffold to operate autonomously under adaptively sampled constraints. It uses a composite reward coupling answer accuracy with budget efficiency. Experiments on seven QA benchmarks show it outperforms baselines across all budget scales and generalizes to unseen constraints.

Original post →

More from coding & agent

coding & agent channel →