IterSynth: Role-Decoupled Search Agent Beats Prior 8B Peers by 4.2%
ZJU-REAL · hf · 2026-09-25
ZJU-REAL proposes IterSynth, a role-decoupled paradigm for deep search agents.
- Problem: ReAct-style agents suffer from role coupling (one policy handles planning, evidence use, and synthesis) and context accumulation noise
- Method: A Planner and a Synthesizer alternate, with an evolving summary as persistent search state; Role-Decoupled Policy Optimization (RDPO) combines outcome rewards with turn-level rubric evaluations and role-specific advantages
- Results: IterSynth-8B averages 50.7 across five long-horizon benchmarks including BrowseComp and Xbench-DS, beating the strongest prior ≤8B agent by +4.2%; as a prompting paradigm it delivers large zero-shot gains over ReAct on frontier proprietary models
More from coding & agent
- Mnemos.Field nears launch: a virtual world where humans and AI agents both register and participate — RileyRalmuto · 2026-09-25
- SemIf open-sources a Jev-style interface: typed option probabilities from a 4B model, no JSON parsing — JeremyCMorgan · 2026-09-25
- Agent Detection-1 launches to tell whether your website visitors are humans or AI agents — IndraVahan · 2026-09-25
- Engineer predicts human-oriented programming languages will die in the AI coding era — kieranklaassen · 2026-09-25
- Agent Arena Ranks 43 Models on 2M+ Real-World Agentic Tasks; Claude Fable 5.1 Tops Board — arena · 2026-09-25
- One prompt: Claude agent wired to Runway MCP delivers a Netflix-style superintelligence doc — CurieuxExplorer · 2026-09-25