Primacy Bias in RLVR Training Paper Accepted to NeurIPS 2026

A NeurIPS 2026 paper by Parshin Shojaee et al. reveals a 'primacy bias' in RLVR training: LLMs disproportionately retain early training samples and struggle to overwrite them, offering new insight into reinforcement learning dynamics.

2026-09-25 ~ 2026-09-25 · 2 related posts