Qwen3.8-Flash Runs Locally in 75GB, GGUF Quantized Versions Released

Unsloth enabled local running of the 125B multimodal MoE model Qwen3.8-Flash in just 75GB of RAM, and released 1-4bit GGUF quantized versions deployable via llama.cpp.

2026-08-26 ~ 2026-08-27 · 2 related posts

Full story(3 episodes)→