Researcher Claims Technical Safety for Closed AI Models is Solved, Blames Incidents on Corporate Shortcuts
StephenLCasper · x · 2026-08-13
AI safety researcher Stephen Casper argues that technical safety for closed-weight AI systems is, at this point, a solved problem. He defines technical safety as the challenge of writing a robust safety specification and getting the model to align its behaviors with it.
He points out that SOTA safety tools are already capable of stopping most technical violations. However, the real-world failures of closed-weight model safety stem from the fact that AI companies take too many shortcuts during deployment. Consequently, he states that he no longer works on safeguards for closed-weight models.
More from AGI Musings
- Academic Shifts View: Under 5% Chance AI Hits a Wall, Expects Full Labor Substitution — aran_nayebi · 2026-08-13
- SPAR Launches Research Project Comparing Animal and AI Welfare — aran_nayebi · 2026-08-13
- Beyond Trivialities: How LLMs Enable Deeper CS Learning — generativist · 2026-08-13
- 'LLM Psychosis': Just Newbies Exploring with High Variance — generativist · 2026-08-13
- Prediction: 2026 Is the Last 'Normal' Year Before Broad Superintelligence — imjustnewatai · 2026-08-13
- xAI Sued by Former Employee Over AI Safety Concerns, Sparking Industry Culture Debate — ryan_t_lowe · 2026-08-13