MIT Tech Review: We're putting too much faith in AI's ability to say no
MIT Tech Review AI · rss · 2026-10-09
MIT Technology Review published an in-depth essay on refusal, the load-bearing wall of modern AI safety.
Key points
- Early models refused nothing — OpenAI's 2022 red-team effort (including Paul Röttger) found the model would happily write Al Qaeda recruitment posts; those datasets were later used for fine-tuning to teach refusal.
- Refusal is not moral reasoning but pattern matching in activation space — recent research describes it as activations forming "high-dimensional polyhedral cones," which can be removed to bypass refusal.
- Refusal is probabilistic and inherently unreliable: determined actors already break through, while models' harmful capabilities scale with their helpful ones — latest models rival top human hackers at network intrusion.
- There is no formula for where to draw the line between obedience and refusal; currently AI companies draw it in secret, and oppressive governments may someday use refusal to stifle legitimate speech.
- The essay warns that failed refusal could bring global calamity, while over-refusal could enable repression — and machines drawing the line themselves would be the worst outcome of all.
More from AGI Musings
- David Patterson predicts SI and robots will rapidly replace jobs by 2029-30 — davidpattersonx · 2026-10-09
- Yacine: courage × intelligence has an optimum — don't be smart enough to think it through — yacineMTB · 2026-10-09
- Researcher slams sweeping claims on consciousness as failing undergrad-level rigor — dioscuri · 2026-10-09
- 2021 Was a Dead Cat Bounce — Only ChatGPT Saved Tech Investing From Ideas Drought — menhguin · 2026-10-09
- Dioscuri: most consciousness scientists back mainstream cognitive theories — dioscuri · 2026-10-09
- Yi Ma and Yann LeCun spar over automation's coming reshaping of mathematics — CSProfKGD · 2026-10-09