Persona selection models falter in high-compute RL, argue AI researchers

voooooogel · x · 2026-09-05

voooooogel responds to BronsonSchoen's critique of the persona selection model (PSM): in high-compute RL regimes PSM becomes an increasingly bad predictor of cognition, and in realistic setups models show behavioral profiles similar to all other reward seekers. voooooogel still finds RL generalization's success strangely lucky, even granting that the NEM result may have been an artifact.

Related event: Persona selection models fail to predict RL outcomes, researchers find(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →