SFTMill: open-source off-policy distillation via any OpenAI-compatible endpoint
jjusko20 · reddit · 2026-09-30
The author open-sourced SFTMill, an off-policy distillation tool that turns any behavioral goal into a full fine-tuning dataset via an OpenAI-compatible endpoint, extracted from his internal tooling for fine-tuning AliceAI 80B A3B. Workflow: define a curriculum in YAML (e.g. tool calls, bug fixing, workspaces, error tracing for an agentic model), an LLM generates tasks (Qwen 3.8 27B medium recommended minimum), then the teacher model solves each task to produce the Q/A set. Repo ships docs and a multi-turn hybrid-reasoning format example.
More from coding & agent
- Denying GPT-6.1 Sol write access as orchestrator cut coding costs 77%, at 6x runtime — GapNew4766 · 2026-09-30
- 8 research agents self-train a 30B model for 144 hours in RSIArena livestream experiment — my_cat_can_code · 2026-09-30
- OpenAI's Nan Yu: a boss agent running other agents is just one agent with extra steps — victor_explore · 2026-09-30
- Open-source MCP server lets agents query 124.6B TikTok data points — operatorarkay · 2026-09-30
- Dev argues truly always-on autonomous agents have never actually been tried — jacob_posel · 2026-09-30
- Skill scaffold template: same shape, faster review, fewer broken CIs — blaizedsouza · 2026-09-30