LFM2.5 training details: MOPD and agentic RL are the two most interesting stages

SergioPaniego · x · 2026-08-04

In LFM2.5-2.6B training, the two most interesting stages: MOPD (student generates, each prompt routes to domain teacher for token-level feedback, teachers branch from same SFT checkpoint, signal close to student distribution) and agentic RL (multi-turn GRPO in real harnesses, one sandbox per rollout, proxy captures token-level trajectories while harness stays black box).

Related event: LFM2.5-2.6B Released: Small Model Beats Larger Counterparts(3 posts)→

Original post →

More from Models

Models channel →