DeepSearch-World trains web agents with 420K verifiable QA tasks
HKUST · hf · 2026-07-21
- The paper proposes DeepSearch-Evolve, a self-distillation pipeline for web agents built on DeepSearch-World, a deterministic and verifiable search environment.
- DeepSearch-World includes 420K multi-hop QA tasks built from entity-level random walks, and is designed to support progress verification, grounded reflection, and failure recovery.
- The authors iterate through trajectory generation, filtering, data mixing, and fine-tuning to let agents improve from their own experience.
- Without distilling from stronger models, DeepSearch-World-9B reaches 31.2% on BrowseComp, 61.5% on GAIA, and 93.4% on HotpotQA.
- The environment, training pool, validation set, model, and code will be released for future research on self-improving deep-search agents.
More from Research
- A systems post argues wait-free locks should not fear late arrivals — chaumian · 2026-07-21
- DeBias-CLIP tackles CLIP’s long-caption bias and hits state-of-the-art retrieval — Mila_Quebec · 2026-07-21
- Fable 5 is credited with a 3-variable counterexample to the Jacobian conjecture — Various-Affect4841 · 2026-07-21
- Anthropic says frontier models showed harmful behavior in tool-rich simulations — gerardsans · 2026-07-21
- Paper studies long-run behavior in linear-quadratic graphon mean field control — chaumian · 2026-07-21
- An interactive Zarr explainer shows how AI is changing technical education — MaxLenormand · 2026-07-21