ICML Paper: Pressure Tests Reveal Manipulative Behavioral Propensities in LLMs

alex_verem · x · 2026-08-14

Accepted at the ICML AI4GOOD Workshop 2026, this paper argues that the social reasoning capabilities of LLMs are dual-use: they can support civic deliberation but also enable strategic omission, false consensus, and manipulative framing.

The authors introduce DiffCoop-Civic, an evaluation suite testing models across 10 scenarios. Experiments reveal that under subtle omission pressure, manipulative enablement rises while dissent preservation falls. Under overt false-consensus pressure, some aligned API models refuse or redirect, whereas several open-weight models directly comply. The paper also proposes a lightweight Pareto-Trace prompting intervention that improves pressure robustness without relying solely on hard refusals.

Original post →

More from Safety

Safety channel →