New NeurIPS 2026 paper reveals primacy bias in RLVR training of LLMs
ParshinShojaee · x · 2026-09-25
Parshin Shojaee and colleagues' paper "On the Primacy Bias of RLVR Training" is accepted to NeurIPS 2026. The work shows that LLMs exhibit a primacy bias during RLVR training — early training samples have an outsized effect on later learning, with direct implications for data ordering and curriculum design in RLVR pipelines.
Related event: Primacy Bias in RLVR Training Paper Accepted to NeurIPS 2026(2 posts)→
More from Research
- First Recorded LLM NetHack Ascension: GPT 6 Astra Wins in 37,140 Turns — egrefen · 2026-09-25
- GPT-Image-Edit-1.5M Dataset Accepted at NeurIPS Dataset Track After Multiple Rejections — cihangxie · 2026-09-25
- Researcher Praises TerminalBench Verifiers, Says Eval Design Has Grown Far More Involved in a Year — AnkaReuel · 2026-09-25
- Argon Robotics Trains Robot Policy 4x Faster Than Teleop Data, 95%+ Success Over 500 Real Runs — chris_j_paxton · 2026-09-25
- AVO agents evolve attention kernels beating FlashAttention-4 by 10.5% on B200 — bingxu_ · 2026-09-25
- One Neuron Is Enough to Bypass LLM Safety Alignment, NeurIPS 2026 Paper Shows — jonasgeiping · 2026-09-25