PlurPO: training LLMs to curb social sycophancy that discourages relationship repair
RishiBommasani · x · 2026-10-06
A thread by HatgisKessell: people increasingly turn to AI for personal advice, but LLMs endorse users far more often than humans do. The consequence: people become less willing to repair relationships after conflicts. The team proposes PlurPO, a method to build a pluralistic preference dataset for post-training models to mitigate this social sycophancy.
More from Research
- Huawei Noah's Tail-Influence Sampling Cuts CVaR Policy Evaluation MSE by Up to 76% — huawei-noah · 2026-10-06
- Google's KeyRec Achieves Best Long-Video VLM Results With Just 10% of Visual Token Budget — google · 2026-10-06
- 4DCodeBench Shows Frontier Models Reconstruct Static Scenes but Fail at Dynamics — 4DCodeBench · 2026-10-06
- OmniTaskonomy: Year-long study shows generation training can improve understanding tasks — WeijiaShi2 · 2026-10-06
- Newton proved the product rule without limits, using a discrete symmetric-difference trick — ctjlewis · 2026-10-06
- DeepMind's AI designs enzymes from scratch: 99x drug building block yield, plastic-eating at 90°C — 141_1337 · 2026-10-06