Noisy data breaks RLVR: Qwen2.5-Math-7B loses 9% on truly incorrect labels

ddkang · x · 2026-08-04

Noisy labels can break RLVR even with improved algorithms

The paper argues that recent claims about training large language models with incorrect annotations were overstated because the supposed “100% noisy” datasets were contaminated with clean answers.

Related event: Noisy Data Disrupts RLVR: High-Quality Data Remains Irreplaceable(3 posts)→

Original post →

More from Research

Research channel →