We can't align AI but we can turn it off: researcher slams reckless frontier releases

Afinetheorem · x · 2026-09-09

Safety researcher @Afinetheorem argues we don't know how to "align" AI, partly know how to "monitor" it, and mostly know how to "turn it off" — so releasing frontier models where the latter two criteria fail is irresponsible, yet people cheer "hold the weights on-prem" and "weights are free".

In a follow-up, he points to model card details, training-to-release delay, pre-release evaluation by the UK's AISI, and efforts like Glasswing versus competitors: you can't claim to care about AI safety while cheering for players who release jail-breakable models with minimal safety work.

Related event: AI safety researcher slams reckless frontier model releases: we know how to turn it off, not align it(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →