R-Quest Fixes Self-Evolving Reasoning Models, Beating R-Zero by 17.32 Points

HINT-lab · hf · 2026-10-08

HINT-lab's paper explains why self-evolving reasoning models collapse over rounds: invalid questions accumulate (and answer-consistency filtering makes it worse), and lexical-similarity diversity control misses mathematically equivalent duplicates. R-Quest trains a solver to reject invalid questions, uses its judgments to reward the questioner and filter training data, and gets novelty feedback via a frozen base model comparing question pairs. It tops 12 benchmarks across two model families, sustains gains over ten self-evolution rounds, and outperforms R-Zero by 17.32 points.

Original post →

More from Research

Research channel →