GGUF quantized Qwen3.8-9B-Distill lands for llama.cpp local runs

empero-ai · hf · 2026-08-19

empero-ai followed up with Qwen3.8-9B-Distill-GGUF, a quantized version for llama.cpp local deployment, with gated-deltanet architecture tags alongside distillation and reasoning. Aimed at on-device use, it's also trending on HF.

Related event: Qwen3.8-9B Distill Model Gets GGUF Release for Local Use(2 posts)→

Original post →

More from Models

Models channel →