Primacy bias exists in RLVR training: new paper accepted to NeurIPS 2026
ParshinShojaee · x · 2026-09-25
A new paper, "On the Primacy Bias in RLVR Training" by Parshin Shojaee et al., has been accepted to NeurIPS 2026. The work shows that LLMs exhibit a primacy bias during RLVR (reinforcement learning with verifiable rewards) training — early exposure leaves lasting preferences that shape later learning. Details on the project page.
Related event: Primacy Bias in RLVR Training Paper Accepted to NeurIPS 2026(2 posts)→
More from Research
- First Recorded LLM NetHack Ascension: GPT 6 Astra Wins in 37,140 Turns — egrefen · 2026-09-25
- GPT-Image-Edit-1.5M Dataset Accepted at NeurIPS Dataset Track After Multiple Rejections — cihangxie · 2026-09-25
- Researcher Praises TerminalBench Verifiers, Says Eval Design Has Grown Far More Involved in a Year — AnkaReuel · 2026-09-25
- Argon Robotics Trains Robot Policy 4x Faster Than Teleop Data, 95%+ Success Over 500 Real Runs — chris_j_paxton · 2026-09-25
- AVO agents evolve attention kernels beating FlashAttention-4 by 10.5% on B200 — bingxu_ · 2026-09-25
- One Neuron Is Enough to Bypass LLM Safety Alignment, NeurIPS 2026 Paper Shows — jonasgeiping · 2026-09-25