ByteShape ships full Qwen 3.8 27B quants: 3.84bpw hits 99.63% of BF16, KLD debunked as quality metric
enrique-byteshape · reddit · 2026-09-15
ByteShape released full ShapeLearn GGUF quants of Qwen 3.8 27B with benchmarks across six GPUs (RTX 6000 Pro Blackwell down to 5060 Ti):
- Accuracy: 3.84bpw (GPU-5) reaches 99.63% of BF16's aggregate score over 8 benchmarks; 3.23bpw still hits 98.72%. All five quants sit on the quality/speed-bpw frontier vs Unsloth v3, ISTA-DASLab, AtomicChat and Bartowski.
- KLD isn't a leaderboard: Unsloth's UD-IQ3S has 20% lower KLD than a similar-size Lite model yet scores worse (95.55% vs 97.33%). KLD only catches broken quants — the thesis of their EMNLP 2026 Industry Track paper.
- Speed: DFlash2 gives 1.34–2.10× throughput; MTP 1.28–1.66× with temperature sampling.
- Benchmarks cover GSM8K, MMLU, IFEval, LiveCodeBench V6, Multi-IF, ACEBench, Multiple HumanEval and BFCL V4 across instruct and thinking modes.
More from Infra
- Hitachi Energy to invest $528 million in new transformer factory in Mississippi — oilmutt · 2026-09-16
- Latham & Watkins, No.2 US Law Firm, Buys Nvidia Hardware to Fine-tune Open Weights In-house — MikeBirdTech · 2026-09-16
- Anthropic, Fluidstack and Cipher pledge $10M to fix a Texas town's water system — MxMnr · 2026-09-16
- Oracle CFO says she 'really, really' dislikes 'doing more with less' a day after layoffs — mkheck · 2026-09-16
- Astra optimizes its own inference on Rubin chips, doubling throughput in 72 hours — bookwormengr · 2026-09-16
- Audio8 open-sources on-device ASR/TTS models down to 0.1B, including iPhone offline transcription — FinanceYF5 · 2026-09-16