Liquid AI Proposes Antidoom Training to Eliminate Model Inference Loops
max_paperclips · x · 2026-08-02
Reasoning models often fall into repetitive 'doom loops' during complex tasks. Liquid AI introduced 'Antidoom', a new method using Final Token Preference Optimization (FTPO) to address this.
The technique identifies the exact token starting a loop and trains the model to prefer coherent alternatives at that position, leaving the rest of the distribution untouched. On the LFM2.5-2.6B model, this reduced loop occurrence in math and coding tasks from 10.2% to 1.4%, with overall evaluation scores improving as a result.
More from Research
- Netflix details its production LLM judge: hundreds of thousands of recommendations scored weekly — omarsar0 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24
- New Architecture RHEA: Train 1B Model on 8GB VRAM — zemondza · 2026-08-24
- Trained two 16M-param models to do generative CAD with real physics — debreuil · 2026-08-24
- Claude model helps discover complex structure on S^6, solving 60-year-old math problem — Singularitarian · 2026-08-24
- Study: Agents read instructions/notes 60.5% of the time, rarely touch API docs — dair_ai · 2026-08-24