HF engineer's DIY continual learning bench: SFT works, SDPO doesn't yet
ben_burtenshaw · x · 2026-09-07
A Hugging Face engineer shares a plan for a personal continual learning benchmark: build RL environments around customizable daily tasks (like day planning), inject synthetic preferences, have a judge model score agent traces, and benchmark LiquidAI/LFM2.5-1.2B-Instruct as the base.
Pipeline: SFT the model on traces, then SDPO with judge-model hints. Status: evals and SFT work; SDPO fails — likely because preferences are too arbitrary, so the author plans to hand-write more preferences for a better few-shot judge.
More from coding & agent
- Design Docs Are All You Need: DeepMind's library regenerates all code from NL docs — omarsar0 · 2026-09-07
- Agents are only powerful if they connect to the systems you already use — shensi · 2026-09-07
- Teknium's agent-driven cleanup sheds 375,000 lines from Hermes Agent codebase — Teknium · 2026-09-07
- Cursor makes self-hosted cloud agents generally available for enterprise networks — thione · 2026-09-07
- Anthropic Open-Sources Claude Commerce Agents, Shopping and Merchant Agent Blueprints — thione · 2026-09-07
- Tactical programming is dead: Kent C. Dodds says fall in love with problem-solving instead — mattpocockuk · 2026-09-07