Noisy Data is Destructive to RLVR: High-Quality Data Remains Essential
ddkang · x · 2026-08-04
A new paper reveals that noisy data is highly destructive to Reinforcement Learning with Verifiable Rewards (RLVR).
- Debunking Prior Claims: Previous studies suggested improved RLVR algorithms allow models to learn from incorrect annotations. This paper shows those datasets were "contaminated" with correct answers due to multiple valid mathematical expressions.
- Rigorous Re-verification: Using GPT-5 Pro and human experts, the authors built a truly 100% noisy dataset. Experiments with Qwen2.5-Math-7B showed an 8-10% accuracy drop compared to training on clean data, performing even worse than format-only rewards.
- Algorithmic Limits: SOTA algorithms like SAPO, DAPO, and DR.GRPO fail to mitigate noise, performing similarly to basic GRPO.
- Capability Degradation: RLVR with noisy data not only lowers accuracy but also weakens reasoning (lower pass@k) and produces increasingly shorter outputs (5-24% reduction).
- Conclusion: Current RLVR methods cannot compensate for poor data quality; high-quality data remains essential for strong reasoning.
Related event: Noisy Data Disrupts RLVR: High-Quality Data Remains Irreplaceable(3 posts)→
More from Research
- One Layer Deeper launches an H100-only competition for deeper reasoning — aryaman2020 · 2026-08-04
- Locus says it beat most human teams across live public ML competitions — rohanpaul_ai · 2026-08-04
- Pure VLAs may not need long-horizon planning if VLMs can cover it — m_wulfmeier · 2026-08-04
- Anthropic says Fable 5 reproduced 5 of OpenAI’s 10 Astra math advances in 24 hours — EricBuess · 2026-08-04
- New CCN poster finds LLM-brain alignment scales differently across cortical systems — neuranna · 2026-08-04
- ThursdAI explores whether Codex and multi-agent math can solve Erdős problems — thursdai_pod · 2026-08-04