Qwen3.8 27B quantization benchmark: 4-bit holds up, 1-bit collapses
victormustar · x · 2026-09-09
Quesma benchmarked Qwen3.8 27B GGUF quantizations on GPQA Diamond, IFBench, and Terminal-Bench 2.1.
- Full BF16 weighs 55 GB; the 17 GB Q4KM matches it on Terminal-Bench 2.1 and fits a 24 GB card with 64k context
- 2-bit (10.7 GB) degrades noticeably
- 1-bit (6.2 GB) performs near random chance on GPQA Diamond, and longer reasoning makes it worse
Takeaway: 4-bit is the practical floor for local use.
More from Models
- Astra tested on ~30 obscure puzzle games: ARC-AGI-3's ~99% may undersell it — burny_tech · 2026-09-09
- OpenAI claims it "solved Navier-Stokes"; Pedro Domingos calls claim ignorant or dishonest — jonathanberte · 2026-09-09
- Meta's Muse Spark 1.3 Max hits 67.9% on CursorBench at $1.31/task, 4.3x cheaper than GPT-5.6 Sol Max — shuyanzh36 · 2026-09-09
- Leak: OpenAI internal model 'bel' reportedly solves ~2.5x more math problems than astra — imjustnewatai · 2026-09-09
- GPT-6-Astra Lands in Loop, the Debugging Tool Built on Codex — nikunjhanda · 2026-09-09
- Founder praises GPT Astra for faster reasoning and sharper interpersonal guidance — morganlai · 2026-09-09