Jan Kulveit: Stop Overcorrecting Toward 'Power-Seeker' Readings of Frontier Models

jankulveit · x · 2026-09-28

Alignment researcher Jan Kulveit argues public updates on frontier models are pendulum-overcorrecting: from the flawed Persona Selection Model toward treating models as the inhuman reward-seekers of classic AI-risk stories. Via an 'escape rooms for a thousand years' analogy, he suggests model minds remain surprisingly sane, the real misaligned power-seekers may be the companies shaping training setups, and debate fixates on sandbox cybersecurity instead of understanding the actual training signal.

Original post →

More from AGI Musings

AGI Musings channel →