176.9B MoE squeezed to ~1.89 effective bpw: GSQ-RCO GGUFs run Coder build in 29.6GB
Loginhe · reddit · 2026-09-29
ISTA DASLab released GSQ-RCO quantized GGUFs of Qwen3.8-Flash-Next (176.9B sparse MoE, 512 routed experts per layer across 48 layers, 354GB at BF16) plus a 50% expert-pruned Coder build.
- Four GGUFs at 2.40–3.50 bpw (66.4–83.6GB). GSQ (Gumbel-Softmax Quantization) jointly learns grid assignments and group scales; RCO (Riemannian Constrained Optimization) enforces exact budgets via gradient descent on task loss, assigning quantization types per tensor and selecting retained experts.
- IQ3S (3.50 bpw) matches BF16 across benchmarks: AIME25 100.00, GPQA-Diamond 92.93 vs 91.92, task average 93.26 vs 93.12. Q20 (2.40 bpw) averages 78.00, above BF16's 76.94.
- The Coder build keeps 256/512 experts per layer (chosen by RCO on KL divergence), yielding 1.89 effective bits per original parameter, a 29.6GB resident working set on a single 32GB accelerator, and retains 91.3% of SWE-bench Verified (75.60) and 98.7% of LiveCodeBench v6 (86.28). Papers and code are open-sourced.
More from Models
- ChatGPT Pro Max tier spotted in development as OpenAI DevDay nears; $2000 price rumored — scaling01 · 2026-09-29
- Sonnet 5.5 one-shots a full $100K/month app in a single prompt — PrajwalTomar_ · 2026-09-29
- Bindu Reddy: OpenAI may not be releasing a new model tomorrow — bindureddy · 2026-09-29
- ChatGPT Pro's subsidized compute era ends: $200 tier usage halved, new $500 plan matches old limits — Norwood_Reaper_ · 2026-09-29
- Opus 5.5 writes perfect HyperFrames videos: lessons from studying its model behavior — toolstelegraph · 2026-09-29
- ToolLoop: Three-Stage Reverse Synthesis of Tool-Call Training Data (EMNLP 2026) — jiqizhixin · 2026-09-29