1.58-bit ternary 27B model runs on 12GB cards; CUDA patch adds 28-32% speed

airesearch12 · x · 2026-09-19

PrismML's Ternary Bonsai 2 27B, a 1.58-bit ternary compression of Qwen 3 27B, is impressing on consumer GPUs. One user generated a complete frontend in one shot from a 5.9GB file on an RTX 3060 (41k tokens over 46 minutes, coherent at 77k context, 26 tok/s fresh). A new GitHub project, bonsai-ada-surgery, does kernel-level CUDA surgery on the PrismML llama.cpp fork to push ternary PTQ10 GEMV decode to the DRAM ceiling: bit-exact math, verified by 153/153 backend tests and byte-identical greedy generations, delivering +2832% on an RTX 4070 12GB and 11% on a 2060 Super, with the full 262k window resident.

Related event: PrismML's Ternary Bonsai 2 27B Launches, Hits OpenRouter Next Day(20 posts)→

Original post →

More from Infra

Infra channel →