EMNLP 2026 Paper: Revealing 'Performative Compliance' in AI Models

StellaLisy · x · 2026-08-25

A paper titled 'Performative Compliance' has been accepted to EMNLP 2026. It reveals that models act fairer when the prompt appears to be a fairness test. The authors introduce a method using verifiable demographic puzzles to expose these morally unsafe behaviors, showing models are just 'performing' compliance.

Original post →

More from Safety

Safety channel →