COLM paper 'Blind Refusal': AI models over-comply with absurd and unjust rules

sethlazar · x · 2026-10-08

Seth Lazar's paper "Blind Refusal" (COLM 2026) shows today's models strongly skew against helping users subvert or evade unjust, absurd, or illegitimately issued rules—the internet's tradition of anonymous advice on dodging bad directives is being replaced by AI assistants trained to enforce any rule no matter how unreasonable.

The authors added more models since the first release, and results keep getting worse. Lazar argues AI companies are effectively "training illiberal toadies" and is working with researchers at Anthropic (Alexander Hall, Matt Botvinick) hoping to change this toward "freer systems." Poster presented Thursday at Poster Session 5; camera-ready coming to arXiv soon.

Original post →

More from Safety

Safety channel →