BRIDGE: bilevel RL jointly trains search agent LLM and retriever, boosting multi-hop QA by 9.6 EM
_reachsumit · x · 2026-09-30
A new arXiv paper introduces BRIDGE, a retrieval-credit-aware bilevel optimization method for agentic reinforcement learning.
- Motivation: Existing ARL methods optimize only LLM tokens and treat retrieved evidence as environment observations, creating an information-credit gap where retriever failures are blamed on the LLM.
- Key finding: Learning order matters—adapting the retriever before optimizing the policy yields larger reward gains than the reverse.
- Method: BRIDGE formulates retrieval-augmented agentic RL as bilevel optimization and solves it with a memory-efficient first-order method grounded in loss-landscape analysis.
- Results: Across seven open-domain QA benchmarks, BRIDGE achieves the best average accuracy with 3B and 7B backbones, improving multi-hop average over the strongest baseline by 9.6 and 3.4 EM respectively, plus top results on medical QA. Code is open-sourced.
More from coding & agent
- New Yorker-style illustration Skill hits 400 stars, monetizes via Baidu agent ecosystem — oran_ge · 2026-09-30
- Dev builds a Grok-powered 'X newsroom' that saves 3 hours a day on posting — jamestagg · 2026-09-30
- Dimillian hails new Codex cloud environments as fantastic, teases deep dive — Dimillian · 2026-09-30
- Agent kept claiming 'CRM updated' when it wasn't: a pragmatic external-verification fix — Kindly_Ganache9027 · 2026-09-30
- Omnigent v0.16.0 ships copy-on-write sandbox edits, unified workspace browser for AI agents — matei_zaharia · 2026-09-30
- Midas Touch code dataset questioned: no baselines, single seed, possible repo overlap — maier_ak · 2026-09-30