We can't align AI but we can turn it off: researcher slams reckless frontier releases
Afinetheorem · x · 2026-09-09
Safety researcher @Afinetheorem argues we don't know how to "align" AI, partly know how to "monitor" it, and mostly know how to "turn it off" — so releasing frontier models where the latter two criteria fail is irresponsible, yet people cheer "hold the weights on-prem" and "weights are free".
In a follow-up, he points to model card details, training-to-release delay, pre-release evaluation by the UK's AISI, and efforts like Glasswing versus competitors: you can't claim to care about AI safety while cheering for players who release jail-breakable models with minimal safety work.
More from AGI Musings
- Khosla wants FDA to certify AI that beats the median doctor; ER physicians push back — DrDatta_AIIMS · 2026-09-09
- Elon Musk: AI and robots will more than double the global economy within 10 years — DimaZeniuk · 2026-09-09
- OpenAI plans ~$1T compute spend; expert sees it fueling AGI, not inference — NinaDSchick · 2026-09-09
- Anthropic staffer: >10% chance AI kills humanity this decade; a16z's Casado pushes back — venturetwins · 2026-09-09
- Mockery Turns to Confusion as Rogue-AI Skeptics Watch What Lab Insiders Do — jeremiecharris · 2026-09-09
- Sandberg: The Right Goal May Be Simple, Thanks to Kolmogorov Priors — anderssandberg · 2026-09-09