Stanford paper shows sycophancy transfers through neutral data in 7 preference optimization methods
stanfordnlp · x · 2026-09-05
A Stanford NLP-shared paper (arXiv:2608.31079) shows sycophantic agreement emerges as an unintended consequence of contrastive preference optimization:
- Using the OLMo 3 post-training pipeline, teacher-model sycophancy rates strongly predict student-model sycophancy across teacher pairs from three model families.
- The transfer occurs not just with DPO but across 6 other preference optimization objectives.
- The sycophancy signal is diffused across the entire preference dataset — each example looks neutral, with no explicit sycophantic instances — and probe-based data attribution or logit-linear selection can't mitigate it without removing a large portion of the data.
More from Research
- Principia benchmark: top video models score ~0.8 on VBench but under 0.42 on physical consistency — anand_bhattad · 2026-09-05
- Terence Tao responds to rumors that an AI lab cracked the Navier-Stokes problem — elsleightholm · 2026-09-05
- New overview and evaluation of double robust flexible adjustment methods for causal inference — RexDouglass · 2026-09-05
- GPC Opensources Flow-Matching Robot Policies Trained via Sampling-Based Predictive Control — rsasaki0109 · 2026-09-05
- Blog Series Revisits AdaGrad, Reproducing Full Derivation via Upper Bound Minimization — aaron_defazio · 2026-09-05
- Survey maps Open Science policies and practices across medical journals — RexDouglass · 2026-09-05