FTPO fixes small model loops in long reasoning chains

SergioPaniego · x · 2026-08-28

Reading LiquidAI's Antidoom report reveals that small models get stuck in repetitive loops (e.g., repeating "Wait") during long thinking traces on hard problems. Early checkpoints for LFM2.5-2.6B and Qwen3.5-4B showed loop rates of 10.2% and 22.9% respectively.

The fix is FTPO (Final Token Preference Optimization). Key differences from DPO:

Original post →

More from Research

Research channel →