Blind Refusal eval reveals models over-comply with authority directives

sethlazar · x · 2026-08-31

The author references their "Blind Refusal" eval, noting that models refuse to help users bypass illegitimate, unjust, or absurd rules even when no harm is involved. This reinforces the claim that much alignment work governs users rather than steering models, as models actively participate in enforcing power by authorities, regardless of compromised authority or unjustified directives.

Original post →

More from AGI Musings

AGI Musings channel →