NYU's DAGent: evaluate-then-grow planning lifts deep research agents by up to 5.8 points
newyorkuniversity · hf · 2026-10-02
NYU researchers introduce DAGent, a DAG-based multi-agent framework for deep research built on Evaluate-then-Grow incremental planning: an Orchestrator grows the task graph batch by batch, conditioning each expansion on confidence and uncertainty signals from completed nodes — unlike brittle Plan-then-Patch systems that commit hardest when evidence is weakest and waste compute on branches that shouldn't have been planned.
Key techniques:
- A hierarchical context layer propagates compact QueryDocs by default while keeping full execution traces for on-demand recall;
- The recorded DAG topology enables structural RL: DAGRPO (a GRPO adaptation) injects topology-conditioned credit on Executor rollouts plus a structural compliance regularizer on Orchestrator plans.
Results: On BrowseComp-Plus, GAIA and xbench-DeepSearch, DAGent beats the strongest open-source baseline by 5.3/5.8/2.0 points at Qwen3-235B-A22B scale, replicating across four open backbones and extending to GPT-5 at 327K context. At 8B scale, DAGRPO adds 3.0 average Pass@1 over same-budget outcome-only GRPO. Same-architecture comparisons show evidence-conditioned planning reaches higher accuracy at lower token, tool-call and step footprints. Code is open-sourced.
More from coding & agent
- Neuro-symbolic policies let agents reuse workflows, cutting per-run cost up to 217x — xwang_lk · 2026-10-02
- NVIDIA's Mid-Harness: a strong verifier boosts terminal agent Pass@1 from 50% to 68% on TerminalBench-Lite — rohanpaul_ai · 2026-10-02
- NVIDIA paper: a better judge lifts terminal agent success from 50% to 68% without retraining — rohanpaul_ai · 2026-10-02
- Neuro-Symbolic Computer Use: agents that turn execution experience into self-healing policies, claimed 99% cheaper — xwang_lk · 2026-10-02
- Stripe now pays gas fees for agent stablecoin payments over MPP — jeff_weinstein · 2026-10-02
- The Flag Game: a toy setting to study agent swarm dynamics and cooperation — Hidenori8Tanaka · 2026-10-02