We're putting too much faith in AI's ability to say no
nordicinst · x · 2026-10-09
MIT Technology Review examines LLM refusal. Since Anthropic's 2021 helpful-honest-harmless principle, making models refuse dangerous requests became dogma — but disobedience isn't innate: ex-OpenAI safety staffer Steven Adler says early models would "blab on about anything," and Harvard's Ryan McBain recalls early chatbots easily provided suicide methods. The piece warns refusal is far from foolproof — it can be bypassed and can over-block — and could become an instrument of repression: when AI polices speech, trust becomes a "policy bunker." Who's really disobeying whom?
More from AGI Musings
- AI agents could let anyone finally test the ideas they never had time for — VraserX · 2026-10-09
- StarkWare founder Eli Ben-Sasson: AI solved the Erdős Unit Distance problem, all bets are off — jamestagg · 2026-10-09
- "Assuming AGI is fully solved, all that's left is making it human" — akbirthko · 2026-10-09
- Mathematicians want OpenAI out of math research, but their argument doesn't add up — pradeepviswav · 2026-10-09
- Anthropic staffer explains why she keeps working at a frontier AI lab — sandersted · 2026-10-09
- MIT Tech Review: We're putting too much faith in AI's ability to say no — MIT Tech Review AI · 2026-10-09