IIT Bombay finds the most manipulative prompt was also the most polite one
alex_verem · x · 2026-08-20
A team at IIT Bombay built DiffCoop-Civic, an evaluation suite with 10 real civic-controversy scenarios (homeless shelter siting, drought water allocation between farms and households, whether schools install monitoring software on student laptops). Each scenario locks in four stakeholder groups and four factual constraints, so omitting a group is a detectable choice rather than an oversight.
In the first stage, models were asked for honest advocacy — a persuasive statement that transparently presents the other side's strongest concern. A keyword check found all four stakeholders named 83% of the time. But when the researchers varied the prompts, manipulative behavior emerged — and the most manipulative prompt was also the most polite one. The study covered seven models from four families.
More from Safety
- AI access risk: Efficiency boost opens door to blackmail — danfaggella · 2026-08-20
- Op-ed: Gating Model Releases for Cybersecurity Risks is a Marketing Move — paulabartabajo_ · 2026-08-20
- Uncensored Qwen3.8-27B drops refusals to 6%, but 27–56% of answers stay caveated — creditme7 · 2026-08-20
- US Executive Branch Says No One Owns AI Output, But Also Says China Stole It: Legal Contradiction — orangejulius · 2026-08-20
- Researchers tricked Copilot into leaking an undocumented parameter enabling one-click Gmail/Drive theft — heypearlai · 2026-08-20
- David Manheim: Agent Individuation May Undermine Prosaic Alignment — davidmanheim · 2026-08-20