Noisy data breaks RLVR: Qwen2.5-Math-7B loses 9% on truly incorrect labels
ddkang · x · 2026-08-04
Noisy labels can break RLVR even with improved algorithms
The paper argues that recent claims about training large language models with incorrect annotations were overstated because the supposed “100% noisy” datasets were contaminated with clean answers.
- The authors re-verified the dataset with GPT-5 Pro and human experts, removing 16% of instances that were accidentally correct.
- They then retrained Qwen2.5-Math-7B with RLVR (GRPO) on truly noisy data.
- Result: performance dropped by about 9% versus clean-data RLVR, and was even worse than using only format-based rewards.
- Across benchmarks, improved RLVR methods like SAPO, DAPO, TIS, DR. GRPO, and PGFC did not fix the problem; under 50% noise, each showed at least one benchmark with more than 5% accuracy degradation.
- The model trained on noisy data also produced shorter outputs and lower pass@k than the clean-trained model.
Related event: Noisy Data Disrupts RLVR: High-Quality Data Remains Irreplaceable(3 posts)→
More from Research
- WaiT for the Signal adds frequency-aware flow matching and cuts sampling compute by 50% — TimDarcet · 2026-08-04
- Quanta: AI is starting to crack legendary Erdős math problems — kylekabasares · 2026-08-04
- Kevin Pratt claims a randomized algorithm breaks the $2^n$ barrier for graph k-coloring — rrwilliams · 2026-08-04
- Jenga boosts LLM serving GPU memory utilization by up to 79.6% on vLLM — AccBalanced · 2026-08-04
- Walden Robotics says humanoids should amplify craftspeople, not replace them — adnothing · 2026-08-04
- XM paper leads one reader to expect compute prices to keep rising — jfischoff · 2026-08-04