Qwen3.8-Flash-Next released in 1-4bit GGUF formats
MaziyarPanahi · x · 2026-08-27
UnslothAI has released GGUF quantized versions of Qwen3.8-Flash-Next, ranging from 1-bit to 4-bit precision. The author suggests using llama.cpp to run these models.
Related event: Qwen3.8-Flash Runs Locally in 75GB, GGUF Quantized Versions Released(2 posts)→
More from Models
- Gemini 3.5 Transcribe launches with 85+ languages and function calling — osanseviero · 2026-08-27
- Hands-on: Gemini 3.7 Flash impresses in frontend tasks — doodlestein · 2026-08-27
- Benchmarking Qwen3.8 27B Quantizations: 4-bit Holds Up, 1-bit Collapses — pmigdal · 2026-08-27
- GLM-5.3-Flash Review: 10% Cost, Pareto Frontier Performance — ArtificialAnlys · 2026-08-27
- Google announces pricing details for Gemini 3.7 Flash — OfficialLoganK · 2026-08-27
- Unsloth releases GGUF quantization of GLM-5.3-Flash model — unsloth · 2026-08-27