Dr. Free drops difficulty rewards, trains self-evolving search agents via evidence gain
_reachsumit · x · 2026-09-29
The paper introduces Dr. Free, the first self-evolving search-agent training framework that removes difficulty-based proposer rewards.
- Problem: Existing data-free self-evolution methods use solver difficulty as a proxy for question quality, which fails to distinguish questions needing cross-passage evidence from shortcut-answerable ones, and measuring difficulty requires costly repeated rollouts per candidate question.
- Method: Dr. Free samples relational chains from a knowledge graph paired with aligned passages to give questions an explicit multi-hop structure. A question earns a positive information-gain reward only when the answer likelihood under full evidence exceeds the maximum likelihood across shortcut contexts.
- Benefit: The reward uses teacher-forced likelihoods, eliminating pass-rate estimation and cutting proposer training time significantly.
More from Research
- Triangle Splatting SLAM: Imperial College's ECCV 2026 dense RGB-D SLAM with on-the-fly mesh extraction — rsasaki0109 · 2026-09-30
- Manifold opens early access: robotics eval platform runs thousands of GPU-parallel rollouts in 30 mins — paigeinsf · 2026-09-30
- 1,000 AI agents discover new CRISPR-like system in virus DNA within 24 hours — CurieuxExplorer · 2026-09-30
- Explaining just 5% of token positions retains nearly all audit success across 4.7M explanations — aisilab · 2026-09-30
- NTU's Persistence Forcing hits FID 1.63 on ImageNet 256 by heterogeneous refinement in pixel-space DiTs — NanyangTechnologicalUniversity · 2026-09-30
- IBM's Q&D trains proactive agents to ask better questions, beating a 15x larger model — ibm · 2026-09-30