9B Distilled Model Halves Token Usage with No Reasoning Drop: Test Shows

GroundbreakingMall54 · reddit · 2026-08-05

An open-source local AI developer tested the DeepSeek-V4-Flash distill of Qwen3.5-9B, the only option fitting in a 7GB VRAM budget. They compared it directly against the base Qwen3.5-9B with identical parameters to isolate the effects of distillation.

Results showed identical correct answers in 6 out of 8 tasks, revealing no measurable reasoning gap. The true difference lay in token consumption: the distilled version significantly shortened outputs across arithmetic, log needle, and tool-calling tasks (total tokens dropped from 8975 to 5480). The author concludes that at the 9B scale, distillation enforces 'output discipline' rather than boosting intelligence, which is highly beneficial for downstream parsing.

Original post →

More from Infra

Infra channel →