Qwen3.8-27B-Ridge-3.7bpw released, shrinking model size to 11.7GB

udmrzn · x · 2026-08-17

Released Qwen3.8-27B-Ridge-3.7bpw, compressing the model size to 11.7 GiB. It features a GDN-aware quantization strategy where state tensors remain Q80 and Gated-DeltaNet mixers use Q4K, preserving native MTP and the full vision tower. It is now supported on llama.cpp, Ollama, and LM Studio.

Original post →

More from Infra

Infra channel →