Qwen 27B IQ3_XXS beats Bonsai ternary PQ2 by 3x on 16GB VRAM local test
Danmoreng · reddit · 2026-09-19
The author compared two locally runnable options on 16GB VRAM: Qwen3.8 27B IQ3XXS (10.18GiB) vs Bonsai ternary PQ2 (6.42GiB), using the same UI-generation tasks on identical hardware.
Quality was close (both 4/4 tasks, 20/20 assertions), but efficiency differed sharply:
- Wall time 8:00 vs 24:09 — Qwen 3.02x faster
- Output tokens 27,197 vs 84,176 — 3.1x fewer (Bonsai is chattier, hence slower)
- Decode 83.59 vs 64.91 tok/s; speculative acceptance 65.22% (MTP) vs 39.76% (modified n-gram)
Caveats: unscientific test, and llama.cpp ran with MTP while Bonsai lacks it, widening the gap. Results and repo are public.
More from Infra
- The rig built to run Emacs and doomscroll X is now worth more than its owner's car — tetsuoai · 2026-09-19
- Apple M4 sustains 10 instructions per cycle, beating most rivals; M5 speedup explained — lemire · 2026-09-19
- Apple M6 bumps cores to 12 with two super cores; CPUs keep improving fast — lemire · 2026-09-19
- Apple M-series chips gained ~50% Geekbench 6 performance over three years — lemire · 2026-09-19
- Inside OpenAI's inference routing: why the proportional controller had to go — AI Engineer · 2026-09-19
- "Normal people can't afford 2x DGX Spark": local AI hardware cost debate — FlolightC · 2026-09-19