Primacy Bias in RLVR Training Paper Accepted to NeurIPS 2026
A NeurIPS 2026 paper by Parshin Shojaee et al. reveals a 'primacy bias' in RLVR training: LLMs disproportionately retain early training samples and struggle to overwrite them, offering new insight into reinforcement learning dynamics.
2026-09-25 ~ 2026-09-25 · 2 related posts
- New NeurIPS 2026 paper reveals primacy bias in RLVR training of LLMs — ParshinShojaee · 2026-09-25
- Primacy bias exists in RLVR training: new paper accepted to NeurIPS 2026 — ParshinShojaee · 2026-09-25