Thinking Machines open-sources Prompt Distillation recipe to bake long system prompts into model weights
TheZachMueller · x · 2026-10-06
- Thinking Machines added a Prompt Distillation (context distillation) recipe to its open-source tinker-cookbook, fine-tuning an LLM to internalize a long, complex prompt into its parameters so it behaves as if the prompt were present, without carrying it at inference time.
- The method has two stages: a teacher model generates responses to queries while conditioned on the target prompt p; a student model is then fine-tuned to produce those responses without ever seeing p.
- Author Zach Mueller notes there's surprisingly little discussion around prompt-baking OSS models to harnesses like Claude Code/Codex and their huge system prompts; this recipe offers a reproducible path to cut per-request context overhead.
More from coding & agent
- Solo dev ships real client work in 36 hours orchestrating Grok Bot with a multi-agent team — alexcovo_eth · 2026-10-06
- Debate: is Codex's harness replaceable with MCP plugins, and will local models win in 5 years? — pvncher · 2026-10-06
- Grok Bot secretly runs a full Debian Linux VM: 8-core Xeon, 15GB RAM, nested virtualization — Thionne_WTZ · 2026-10-06
- FlowBank (NeurIPS 2026): precomputed workflow portfolios give agents query-level adaptivity at task-level cost — furongh · 2026-10-06
- TagScribeR rebuilt: free local dataset studio with native LoRA training on AMD ROCm and NVIDIA — ArchAngelAries · 2026-10-06
- Dev vibe codes a browser CS2 remake in one week, runs smooth on weak GPUs — TAbrodi · 2026-10-06