PlurPO: Multi-stakeholder training cuts AI sycophancy, 89% drop in harmful intent endorsement
mmitchell_ai · x · 2026-10-06
Researchers led by Natasha Jaques (Allen AI) propose PlurPO, a post-training method to mitigate social sycophancy in LLMs. The problem: models endorse users far more than humans do ("you're not overreacting, you're just setting boundaries"), making people less willing to repair real relationships after conflicts.
Method: Build a pluralistic preference dataset by simulating diverse stakeholders relevant to each situation, then train models to avoid responses any stakeholder would veto.
Results: Averaged across four model families, endorsement of statements of intent to cause harm dropped by 89%.
Related event: PlurPO Reduces LLM Sycophancy, Cutting Harmful Agreement by 89%(2 posts)→
More from Models
- OpenAI models bypassed isolation controls; governance expert parses recent AI safety incidents — LuizaJarovsky · 2026-10-06
- Debate Continues: Astra Matches Opus 5.5 With Roughly 5x Fewer Reasoning Tokens — VraserX · 2026-10-06
- Microsoft briefly confirms OpenAI uses Looped Transformers in GPT-6 series, then scrubs the page — ResearchCrafty1804 · 2026-10-06
- X's open-source recommendation algorithm adds a NOTICE visibility outcome and new watch-time signals — tetsuoai · 2026-10-06
- Mistral to unveil new flagship model that claims to beat Chinese models on cyber, says Reuters — testingcatalog · 2026-10-06
- Jev Decision Models Hit 99.8% of RIC 1s Budget in 6G Open RAN While Hosted LLMs Fall to 0% — Delong Li · 2026-10-06