Inside LFM2.5-2.6B's post-training recipe: SFT, RL, multi-domain distillation

helloiamleonie · x · 2026-09-25

Leonie assembles the full post-training recipe for LFM2.5-2.6B: (1) classical SFT with agent traces and antidoom training; (2) SFT+RL to train specialized teachers per skill; (3) MOPD distills all skills back into one checkpoint; (4) agentic RL polishes the model for agent harnesses. The 2.6B model competes with the 3x larger Qwen3.5-9B.

Related event: Liquid AI Releases LFM2.5-2.6B and Details Its On-Device Training Recipe(10 posts)→

Original post →

More from Models

Models channel →