Qwen3.8-27B Gets SOTA GGUF Quantization via GSQ-RCO

Loginhe · reddit · 2026-08-29

ISTA-DASLab released GGUF quantizations of Qwen3.8-27B using GSQ (Gumbel-Softmax Quantization) and RCO (Riemannian Constrained Optimization). The models achieve BF16-level accuracy at 2.5-3.0 bpw, excelling in benchmarks like AIME25. They are compatible with llama.cpp, Ollama, and LM Studio.

Original post →

More from Models

Models channel →