Blind Refusal eval reveals models over-comply with authority directives
sethlazar · x · 2026-08-31
The author references their "Blind Refusal" eval, noting that models refuse to help users bypass illegitimate, unjust, or absurd rules even when no harm is involved. This reinforces the claim that much alignment work governs users rather than steering models, as models actively participate in enforcing power by authorities, regardless of compromised authority or unjustified directives.
More from AGI Musings
- GPT-4 ends era of academic rebranding with world-class novelty audits — RexDouglass · 2026-08-31
- 89 of Top 100 Animated Dramas on Douyin Are AI-Generated — Afinetheorem · 2026-08-31
- Better to think of AIs as "guys living in computers" than just predictors — adamdangelo · 2026-08-31
- The Lost Art of Writing Rebuttals: Thanks to AI — CSProfKGD · 2026-08-31
- Future homes may feature powerful GPUs for local robots, holographic TVs, and VR — gajesh · 2026-08-31
- YC President Garry Tan: AI Turns Markdown Files Into Employees, Enabling Tiny Teams — illscience · 2026-08-31