EviBack uses teacher backoff to train agentic RAG on all-zero reward rollouts

_reachsumit · x · 2026-07-28

EviBack proposes an evidence-constrained Teacher backoff for agentic RAG training. The method gives auxiliary supervision to rollout groups with all-zero rewards, while keeping answer refinement separate from evidence assessment so reference answers do not mask evidence insufficiency. An automated GPT-5.5-assisted APE pipeline builds a gated two-stage Teacher from a manually written dual-task prompt, and the authors report higher downstream F1 and valid-answer rates with fewer searches, duplicate queries, and forced terminations across seven open-domain QA benchmarks and three Qwen3 scales.

Original post →

More from coding & agent

coding & agent channel →