Qwen 27B IQ3_XXS beats Bonsai ternary PQ2 by 3x on 16GB VRAM local test

Danmoreng · reddit · 2026-09-19

The author compared two locally runnable options on 16GB VRAM: Qwen3.8 27B IQ3XXS (10.18GiB) vs Bonsai ternary PQ2 (6.42GiB), using the same UI-generation tasks on identical hardware.

Quality was close (both 4/4 tasks, 20/20 assertions), but efficiency differed sharply:

Caveats: unscientific test, and llama.cpp ran with MTP while Bonsai lacks it, widening the gap. Results and repo are public.

Original post →

More from Infra

Infra channel →