R-Quest Fixes Self-Evolving Reasoning Models, Beating R-Zero by 17.32 Points
HINT-lab · hf · 2026-10-08
HINT-lab's paper explains why self-evolving reasoning models collapse over rounds: invalid questions accumulate (and answer-consistency filtering makes it worse), and lexical-similarity diversity control misses mathematically equivalent duplicates. R-Quest trains a solver to reject invalid questions, uses its judgments to reward the questioner and filter training data, and gets novelty feedback via a frozen base model comparing question pairs. It tops 12 benchmarks across two model families, sustains gains over ten self-evolution rounds, and outperforms R-Zero by 17.32 points.
More from Research
- Ethereum's Justin Drake calls for 'bunker mode' crypto migration as AI math advances threaten assumptions — marcvanderchijs · 2026-10-08
- Lean proof may not map 1-to-1 to paper: GPT formalization takes shortcuts on hard lemmas — ctjlewis · 2026-10-08
- 'Alignment Whack-a-Mole' gets COLM 2026 oral slot, TV interview teased — TuhinChakr · 2026-10-08
- OpenAI's claimed modularity proof could prove all elliptic curves over CM fields are modular — soumitrashukla9 · 2026-10-08
- OpenAI Published 722 Math Papers in One Day — and Cites Itself in Over Half of Them — aran_nayebi · 2026-10-08
- AI Reveals Insights Hide in the 'Convex Hull' of Existing Ideas — and Science at Large Is Next — soumitrashukla9 · 2026-10-08