Distilling an Agent Down to 2B for Consumer GPUs via DAgger-Style Loops
Kangwook_Lee · x · 2026-09-29
The author shares a full pipeline for shrinking a model to run on consumer GPUs:
- Following the spirit of DAgger: collect trajectories, correct them, then keep training the model.
- After training, deploy the new model for the next round of PC-bang data collection, repeating the loop.
- Finally distill (SFT, off-policy KD, then agentic OPD) the model all the way down to a 2B LLM so it runs on consumer GPUs.
The key is a deploy-collect-correct-retrain closed loop rather than one-shot offline distillation.
More from coding & agent
- StepFun co-founder proposes KITE: PD-separation-inspired training for scaling agentic LLMs — teortaxesTex · 2026-09-29
- LangChain team talk by Sydney Runkle and Victor Moreira is worth watching even if you don't use LangChain — js_craft_hq · 2026-09-29
- CS student asks how shipped agents handle confident-but-wrong actions on real systems — Professional-Mine681 · 2026-09-29
- WebBrain: open-source local browser agent with a 450M browser-specialized VLM — ButtercupLyn100 · 2026-09-29
- NVIDIA OpenShell tested: 10/10 secret leaks without it, 0/10 with default policy — but auto-approve leaked in 12/12 — No-Peanut-6988 · 2026-09-29
- Sonnet 5.5 one-shots a full $100K/month app in a single prompt — PrajwalTomar_ · 2026-09-29