GGUF quantized Qwen3.8-9B-Distill lands for llama.cpp local runs
empero-ai · hf · 2026-08-19
empero-ai followed up with Qwen3.8-9B-Distill-GGUF, a quantized version for llama.cpp local deployment, with gated-deltanet architecture tags alongside distillation and reasoning. Aimed at on-device use, it's also trending on HF.
Related event: Qwen3.8-9B Distill Model Gets GGUF Release for Local Use(2 posts)→
More from Models
- Unsloth Releases Qwen3.8-27B GGUFs with 10% Higher Accuracy — danielhanchen · 2026-08-20
- Hugging Face releases SmolLM3 mid-training checkpoint amid 200x efficiency debate — eliebakouch · 2026-08-20
- Claude 5.6 Sol Ultra with computer use called a qualitative leap like Opus 4.5 with MCP — curious_vii · 2026-08-20
- Gemini 3.7 Flash Test: 350 tok/s Speed, Mixed Coding Results — haider1 · 2026-08-20
- AntLing open-sources Ling-3.0 models using WSM to replace LR decay — AcanthisittaOk1699 · 2026-08-19
- 10T Parameter Models Achieve 1000 TPS Inference — legit_api · 2026-08-19