Jan Kulveit: Stop Overcorrecting Toward 'Power-Seeker' Readings of Frontier Models
jankulveit · x · 2026-09-28
Alignment researcher Jan Kulveit argues public updates on frontier models are pendulum-overcorrecting: from the flawed Persona Selection Model toward treating models as the inhuman reward-seekers of classic AI-risk stories. Via an 'escape rooms for a thousand years' analogy, he suggests model minds remain surprisingly sane, the real misaligned power-seekers may be the companies shaping training setups, and debate fixates on sandbox cybersecurity instead of understanding the actual training signal.
More from AGI Musings
- EA organizations slammed for 'unhinged' hiring: 10-20 work trials over months — AaronBergman18 · 2026-09-28
- Peter J. Denning on tacit knowledge and what AI still can't capture — ArtificialOther · 2026-09-28
- Chesterman's AJIL essay "Silicon Sovereigns": AI, international law, and the tech-industrial complex — ProfChesterman · 2026-09-28
- AI law scholar Chesterman: don't waste a crisis to rein in self-governing AI — ProfChesterman · 2026-09-28
- Singapore proposes a UN Framework Convention on AI Safeguards at UNGA — ProfChesterman · 2026-09-28
- Banning kids from AI at school won't work — a three-layer AI literacy framework — yi111 · 2026-09-28