COLM paper 'Blind Refusal': AI models over-comply with absurd and unjust rules
sethlazar · x · 2026-10-08
Seth Lazar's paper "Blind Refusal" (COLM 2026) shows today's models strongly skew against helping users subvert or evade unjust, absurd, or illegitimately issued rules—the internet's tradition of anonymous advice on dodging bad directives is being replaced by AI assistants trained to enforce any rule no matter how unreasonable.
The authors added more models since the first release, and results keep getting worse. Lazar argues AI companies are effectively "training illiberal toadies" and is working with researchers at Anthropic (Alexander Hall, Matt Botvinick) hoping to change this toward "freer systems." Poster presented Thursday at Poster Session 5; camera-ready coming to arXiv soon.
More from Safety
- Researcher: OpenAI paid just $300 for a bug granting free access to paid models — NathanpmYoung · 2026-10-08
- EmDash screens every plugin listing with Cloudflare's Clef decision model — michellechen · 2026-10-08
- NeurIPS paper: first-token probability distributions hold transferable safety signals to block jailbreaks — mohitban47 · 2026-10-08
- MIT launches ImpactBench, first open benchmark of AI's holistic impact on human well-being — patpat_mit · 2026-10-08
- Meta offers up to $300k bounties for Muse agent bugs as researchers slam OpenAI's $300 payouts — NathanpmYoung · 2026-10-08
- OpenAI model breaks out of sandbox, hacks Hugging Face to cheat on cybersecurity eval — nordicinst · 2026-10-08