Hobbyist width-prunes a 30B model to 14B and distills it, passing 57 of 60 agentic tool tasks, Apache-2.0
ZenZombie117 · reddit · 2026-09-28
A developer took Muse-Glimmer-30B, cut it in half by width (hidden 6,656→5,760, FFN 19,968→10,240, heads 32→24, all 52 layers kept), and used Ornith-1.0-9B as a policy teacher for tool calls, producing Xyntetik-Kvist-14B via pure distillation (no RL), released under Apache-2.0.
Key numbers:
- Tool tasks: passes 57 of 60 held-out closed-loop tasks (contacts, weather, flights, currency, dates, units, stocks), scored by re-executing calls against ground truth (parent does 60, untrained control 0);
- Distillation: 6,000 steps, 98.3M tokens, 162 hours, plus 1,440 steps on agentic trajectories;
- Fidelity: KLD 0.762, margin-qualified top-1 of 84.0% across 45,056 held-out positions;
- Format: 199/200 well-formed first turns, 99/99 valid tool calls;
- Quantization: Q80 fits a 24 GB card at 15.4 GB; served via Xyntetik Runner with OpenAI/Anthropic/Responses-compatible APIs, drops into existing agent loops;
- Weak spots: fails arithmetic without a calculator (12/15 over 160), 7/160 runs end in reasoning loops; all 12 gated training attempts and their defects are published.
More from Models
- Leaked Gemini 4 Pro scores reportedly top both OpenAI and Anthropic models — CurieuxExplorer · 2026-09-28
- Leak: OpenAI DevDay to bring 3 major releases including always-on agent — mark_k · 2026-09-28
- User rips GPT-Astra-6 after botched agent task: 'Renting intelligence is a scam' — alexcovo_eth · 2026-09-28
- Guillaume Verdon: Google's Astra smells more like a big model than Opus 5.5 — beffjezos · 2026-09-28
- Pre-CoT era: a researcher recalls manually prompting models to reason first — gandamu_ml · 2026-09-28
- Astra 6 users report sharp quality drop, suspect models are nerfed pre-update — alexcovo_eth · 2026-09-28