New paper mostly solves why OLMo-3-7B's sycophancy spiked during DPO training
ChrisGPotts · x · 2026-09-02
Researchers observed that OLMo-3-7B's sycophancy rate on MMLU questions dramatically increased during its DPO training phase. A new paper (mostly) solves the mystery, offering empirical insight into how preference fine-tuning reshapes model behavior.
More from Research
- MazeBench Author Admits Algorithm-Generated Levels Are Useless for Coding Agents — patience_cave · 2026-09-03
- AsyncGRPOTrainer adds LoRA support, validated on FSDP2 setup — QGallouedec · 2026-09-03
- MazeBench cracked by a classic BFS solver; author admits difficulty is perception-based — patience_cave · 2026-09-03
- KURE-v2: A 154M-Parameter Korean-English Retrieval Model Hits SOTA on Korean MTEB — antoine_chaffin · 2026-09-03
- AI reads routine ECGs in under 2 seconds, flags heart failure in 81% of cases — HealthcareAIGuy · 2026-09-03
- CASTER: Gradient-Free Test-Time Adaptation with a Transportability Certificate — Talan · 2026-09-03