HF engineer's DIY continual learning bench: SFT works, SDPO doesn't yet

ben_burtenshaw · x · 2026-09-07

A Hugging Face engineer shares a plan for a personal continual learning benchmark: build RL environments around customizable daily tasks (like day planning), inject synthetic preferences, have a judge model score agent traces, and benchmark LiquidAI/LFM2.5-1.2B-Instruct as the base.

Pipeline: SFT the model on traces, then SDPO with judge-model hints. Status: evals and SFT work; SDPO fails — likely because preferences are too arbitrary, so the author plans to hand-write more preferences for a better few-shot judge.

Original post →

More from coding & agent

coding & agent channel →