9B Distilled Model Halves Token Usage with No Reasoning Drop: Test Shows
GroundbreakingMall54 · reddit · 2026-08-05
An open-source local AI developer tested the DeepSeek-V4-Flash distill of Qwen3.5-9B, the only option fitting in a 7GB VRAM budget. They compared it directly against the base Qwen3.5-9B with identical parameters to isolate the effects of distillation.
Results showed identical correct answers in 6 out of 8 tasks, revealing no measurable reasoning gap. The true difference lay in token consumption: the distilled version significantly shortened outputs across arithmetic, log needle, and tool-calling tasks (total tokens dropped from 8975 to 5480). The author concludes that at the 9B scale, distillation enforces 'output discipline' rather than boosting intelligence, which is highly beneficial for downstream parsing.
More from Infra
- Building a Trusted Compute Cluster: Infrastructure for Safe Frontier AI Evaluation — ohlennart · 2026-08-06
- Running MiniMax H3 on Dual RTX 5060 Ti: VRAM Splitting and Unloading Strategies — Kahvana · 2026-08-06
- Running MiniMax H3 on AMD GPUs Hits a Wall with SageAttention Dependency — itiswhatitiswgatitis · 2026-08-06
- Teaser: Run Any Model via One API Key with Local-Cloud Routing — ycombinator · 2026-08-06
- Inside turbopuffer: How to Build a 256TB Search Index — shakoistsLog · 2026-08-06
- GlobalFoundries Q3 Margin Guides Above 30% as Silicon Photonics Capacity Becomes Scarce — tengyanAI · 2026-08-06