Liquid AI's "Antidoom Training" cuts model doom loops from 10% to 1%
helloiamleonie · x · 2026-09-25
A quick refresher on RL with verifiable rewards (RLVR): extract the final answer from the model's response, verify it against ground truth, and grade the trajectory accordingly. This works well but the reward signal is sparse, making long multi-turn generations inefficient.
To address this, the team developed antidoom training: detect a doom loop, find the first token that triggers it, reject and downsample it, and upsample alternative tokens. On LFM2.5-2.6B this reduced doom loops from 10% to 1%.
Related event: Liquid AI Releases LFM2.5-2.6B and Details Its On-Device Training Recipe(10 posts)→
More from Models
- Google criticized for locking $20 Workspace AI subscribers to outdated Gemini models — thedealdirector · 2026-09-25
- Google's $20 Workspace AI add-on slammed for leaving enterprise Gemini stuck on old Flash model — thedealdirector · 2026-09-25
- Hands-on with StepFun Step-5-Preview: rock-solid agent loops, weak on 3D and aesthetics — karminski3 · 2026-09-25
- Claude Opus 5.5 scores 31.2% on WeirdML v3, trailing GPT 6 Astra's 42.2% — scaling01 · 2026-09-25
- Sora API shuts down today with no replacement from OpenAI — VraserX · 2026-09-25
- 'LLMs will eventually write assembly' — Greg Mushen finally saw an example that convinced him — gregmushen · 2026-09-25