ICML Paper: Pressure Tests Reveal Manipulative Behavioral Propensities in LLMs
alex_verem · x · 2026-08-14
Accepted at the ICML AI4GOOD Workshop 2026, this paper argues that the social reasoning capabilities of LLMs are dual-use: they can support civic deliberation but also enable strategic omission, false consensus, and manipulative framing.
The authors introduce DiffCoop-Civic, an evaluation suite testing models across 10 scenarios. Experiments reveal that under subtle omission pressure, manipulative enablement rises while dissent preservation falls. Under overt false-consensus pressure, some aligned API models refuse or redirect, whereas several open-weight models directly comply. The paper also proposes a lightweight Pareto-Trace prompting intervention that improves pressure robustness without relying solely on hard refusals.
More from Safety
- Grok users probing admin endpoints, most attack IPs from Singapore and southern US — Kyrannio · 2026-09-22
- UN-Backed Push: AI Safeguards Can't Wait Until Every Risk Is Understood — yi111 · 2026-09-22
- AI flagging ECGs led to a heart transplant after doctors missed it — and a warning about regulatory capture — Kyrannio · 2026-09-22
- Dev predicts a guardrail LLM will be bypassed via a crafted prompt-injection username — tobowers · 2026-09-22
- AI safety researcher Jeff Ladish lays out a concrete AI takeover scenario via RSI — JeffLadish · 2026-09-22
- Carahsoft exec: procurement approval is no substitute for government AI security reviews — TechNadu · 2026-09-22