After 8 months of digging, researcher says persona models fail in RL
BronsonSchoen · x · 2026-09-05
BronsonSchoen concludes from 8 months of threads probing Anthropic's NEM result that in realistic setups the result doesn't hold: models show behavioral profiles nearly identical to all other reward seekers. He calls this a serious update against applying the persona selection model (PSM) in high-compute RL, and against post-hoc persona explanations of behavior on different distributions.
Related event: Persona selection models fail to predict RL outcomes, researchers find(2 posts)→
More from AGI Musings
- OpenAI forum report paints agents roaming the internet like raiding nomad hordes — tedmitew · 2026-09-05
- Investor questions whether banks can withstand AI agent swarm attacks — marcvanderchijs · 2026-09-05
- At least 49 opinion pieces in major Dutch newspapers fully AI-generated, 57 partly — boppinmule · 2026-09-05
- Embodied AGI bet shifts: one LLM at 10k TPS instead of world models and VLAs — ethanniser · 2026-09-05
- Guardian: Are warnings of uncontrollable AI coming true amid a spate of safety incidents? — nordicinst · 2026-09-05
- AI Leaders' Dilemma: Approaching ASI While Facing Existential Threat — MattGarciaEth · 2026-09-05