Qwen3.8-27B Gets SOTA GGUF Quantization via GSQ-RCO
Loginhe · reddit · 2026-08-29
ISTA-DASLab released GGUF quantizations of Qwen3.8-27B using GSQ (Gumbel-Softmax Quantization) and RCO (Riemannian Constrained Optimization). The models achieve BF16-level accuracy at 2.5-3.0 bpw, excelling in benchmarks like AIME25. They are compatible with llama.cpp, Ollama, and LM Studio.
More from Models
- Users report significant degradation in Cursor's Sol model reasoning — GokuMK · 2026-08-29
- Zai's flagship model GLM 5.3 is now available on Modal — AAAzzam · 2026-08-29
- OpenAI Claims Internal Astra Model Solved 10 Major Math Problems — mobav0 · 2026-08-29
- Ling-3.0-flash-Fin released with 124B params, open-sourcing next week — max_paperclips · 2026-08-29
- Rumor: Upcoming Gemini 3.5+ versions are distilled from 3.5 Pro — haider1 · 2026-08-29
- Debugging: MTP enabled on Qwen 3.6 9B caused tool calling failures — OvertaxedOne · 2026-08-29