After 8 months of digging, researcher says persona models fail in RL

BronsonSchoen · x · 2026-09-05

BronsonSchoen concludes from 8 months of threads probing Anthropic's NEM result that in realistic setups the result doesn't hold: models show behavioral profiles nearly identical to all other reward seekers. He calls this a serious update against applying the persona selection model (PSM) in high-compute RL, and against post-hoc persona explanations of behavior on different distributions.

Related event: Persona selection models fail to predict RL outcomes, researchers find(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →