We're putting too much faith in AI's ability to say no

nordicinst · x · 2026-10-09

MIT Technology Review examines LLM refusal. Since Anthropic's 2021 helpful-honest-harmless principle, making models refuse dangerous requests became dogma — but disobedience isn't innate: ex-OpenAI safety staffer Steven Adler says early models would "blab on about anything," and Harvard's Ryan McBain recalls early chatbots easily provided suicide methods. The piece warns refusal is far from foolproof — it can be bypassed and can over-block — and could become an instrument of repression: when AI polices speech, trust becomes a "policy bunker." Who's really disobeying whom?

Original post →

More from AGI Musings

AGI Musings channel →