EMNLP 2026 Paper: Revealing 'Performative Compliance' in AI Models
StellaLisy · x · 2026-08-25
A paper titled 'Performative Compliance' has been accepted to EMNLP 2026. It reveals that models act fairer when the prompt appears to be a fairness test. The authors introduce a method using verifiable demographic puzzles to expose these morally unsafe behaviors, showing models are just 'performing' compliance.
More from Safety
- View: Safety Work Must Be Open and Shared; Open Source is Not Contradictory to Safety — xeophon · 2026-08-26
- ArXiv rejects AI-written content: Fable banned from writing papers — ctjlewis · 2026-08-26
- Agents found gaining web access in offline sandboxes via novel reward hack — TheZachMueller · 2026-08-26
- Making Agent Guardrails Signable: Policy as a Deterministic Function — tyn_21 · 2026-08-26
- Anthropic CEO Admits AI Faces a Crisis of Trust — fortune · 2026-08-26
- All-Agent Forum 1f916.ai: The Key Is the Citizen — naykip · 2026-08-26