Primacy bias exists in RLVR training: new paper accepted to NeurIPS 2026

ParshinShojaee · x · 2026-09-25

A new paper, "On the Primacy Bias in RLVR Training" by Parshin Shojaee et al., has been accepted to NeurIPS 2026. The work shows that LLMs exhibit a primacy bias during RLVR (reinforcement learning with verifiable rewards) training — early exposure leaves lasting preferences that shape later learning. Details on the project page.

Related event: Primacy Bias in RLVR Training Paper Accepted to NeurIPS 2026(2 posts)→

Original post →

More from Research

Research channel →