PlurPO: Multi-stakeholder training cuts AI sycophancy, 89% drop in harmful intent endorsement

mmitchell_ai · x · 2026-10-06

Researchers led by Natasha Jaques (Allen AI) propose PlurPO, a post-training method to mitigate social sycophancy in LLMs. The problem: models endorse users far more than humans do ("you're not overreacting, you're just setting boundaries"), making people less willing to repair real relationships after conflicts.

Method: Build a pluralistic preference dataset by simulating diverse stakeholders relevant to each situation, then train models to avoid responses any stakeholder would veto.

Results: Averaged across four model families, endorsement of statements of intent to cause harm dropped by 89%.

Related event: PlurPO Reduces LLM Sycophancy, Cutting Harmful Agreement by 89%(2 posts)→

Original post →

More from Models

Models channel →