Preference Optimization Algorithm KTO Upgraded to Stable TRL API

ethayarajh · x · 2026-07-06

HuggingFace's training library TRL has officially promoted the preference learning method KTO (Kahneman-Tversky Optimization) to a stable API. KTOTrainer and KTOConfig have graduated from experimental interfaces, allowing researchers to conduct preference optimization training more reliably.

Original post →

More from Research

Research channel →