Burkov: DPO, IPO, KTO, SimPO, ORPO Unified as Three-Axis Choices in One Framework
burkov · x · 2026-09-15
ML author Andriy Burkov highlights an article that unifies the growing zoo of preference learning techniques behind a single theoretical lens, addressing confusion after the shift from RLHF to simpler direct methods.
- It shows DPO, IPO, KTO, SimPO and ORPO are not unrelated inventions but different choices along three axes: the preference model turning human judgments into an objective, the regularization limiting deviation from a reference policy, and the data distribution governing online vs offline learning.
- The argument draws on formal theorems, coverage arguments, and a synthesis of findings from over fifty papers.
More from Research
- RewardAI Unveils OM-1 Robot Foundation Model: Zero-Shot Generalization From Human Data Alone — yifengzhu_ut · 2026-09-15
- Brain implant lets paralyzed woman converse in real time via AI-generated voice — Polymarket · 2026-09-15
- Could AI training work like Bitcoin? A compute-voting thought experiment — dbasch · 2026-09-15
- Yandex open-sources its search AI answer model, squeezing 40% more answers from same compute — teortaxesTex · 2026-09-15
- AI face-preference study goes megaviral with 450,000 completers, v2.0 released — SpencrGreenberg · 2026-09-15
- Scholars blast arXiv for losing its way as a preprint server over AI crackdowns — RexDouglass · 2026-09-15