Qwen3.8-27B-Ridge-3.7bpw released, shrinking model size to 11.7GB
udmrzn · x · 2026-08-17
Released Qwen3.8-27B-Ridge-3.7bpw, compressing the model size to 11.7 GiB. It features a GDN-aware quantization strategy where state tensors remain Q80 and Gated-DeltaNet mixers use Q4K, preserving native MTP and the full vision tower. It is now supported on llama.cpp, Ollama, and LM Studio.
More from Infra
- Score Studio enables local inference with decentralized storage — markjeffrey · 2026-08-17
- Nvidia to Invest Up to $3B in Lancium, Owner of Stargate Campus Power Infra — Beth_Kindig · 2026-08-17
- Bun's Web APIs get up to 4x faster in next version — ctjlewis · 2026-08-17
- GLM 5.3 Coming to AI Gateway: Top Score in DeepsecBench — evilrabbit_ · 2026-08-17
- Red Hat: Building a production-grade operational layer for AI agents — blaizedsouza · 2026-08-17
- Day 81: Building distributed AI pipelines with observability — blaizedsouza · 2026-08-17