SFTMill: open-source off-policy distillation via any OpenAI-compatible endpoint

jjusko20 · reddit · 2026-09-30

The author open-sourced SFTMill, an off-policy distillation tool that turns any behavioral goal into a full fine-tuning dataset via an OpenAI-compatible endpoint, extracted from his internal tooling for fine-tuning AliceAI 80B A3B. Workflow: define a curriculum in YAML (e.g. tool calls, bug fixing, workspaces, error tracing for an agentic model), an LLM generates tasks (Qwen 3.8 27B medium recommended minimum), then the teacher model solves each task to produce the Q/A set. Repo ships docs and a multi-turn hybrid-reasoning format example.

Original post →

More from coding & agent

coding & agent channel →