Liquid AI's "Antidoom Training" cuts model doom loops from 10% to 1%

helloiamleonie · x · 2026-09-25

A quick refresher on RL with verifiable rewards (RLVR): extract the final answer from the model's response, verify it against ground truth, and grade the trajectory accordingly. This works well but the reward signal is sparse, making long multi-turn generations inefficient.

To address this, the team developed antidoom training: detect a doom loop, find the first token that triggers it, reject and downsample it, and upsample alternative tokens. On LFM2.5-2.6B this reduced doom loops from 10% to 1%.

Related event: Liquid AI Releases LFM2.5-2.6B and Details Its On-Device Training Recipe(10 posts)→

Original post →

More from Models

Models channel →