1.58-bit ternary 27B model runs on 12GB cards; CUDA patch adds 28-32% speed
airesearch12 · x · 2026-09-19
PrismML's Ternary Bonsai 2 27B, a 1.58-bit ternary compression of Qwen 3 27B, is impressing on consumer GPUs. One user generated a complete frontend in one shot from a 5.9GB file on an RTX 3060 (41k tokens over 46 minutes, coherent at 77k context, 26 tok/s fresh). A new GitHub project, bonsai-ada-surgery, does kernel-level CUDA surgery on the PrismML llama.cpp fork to push ternary PTQ10 GEMV decode to the DRAM ceiling: bit-exact math, verified by 153/153 backend tests and byte-identical greedy generations, delivering +2832% on an RTX 4070 12GB and 11% on a 2060 Super, with the full 262k window resident.
Related event: PrismML's Ternary Bonsai 2 27B Launches, Hits OpenRouter Next Day(20 posts)→
More from Infra
- Open models hit record 78.4% of token volume on Vercel AI Gateway; Kimi, DeepSeek, Z.ai spend beats OpenAI — charles_irl · 2026-09-19
- Awesome MCP Servers: a 95k-star index of every production MCP server — mdancho84 · 2026-09-19
- X open-source algorithm update: weights unchanged, video checks start at 64 likes — ThePeterMick · 2026-09-19
- 4x RTX 3090 advice: keep Qwen 27B Q8 or switch to Qwen Next Flash — Zyj · 2026-09-19
- antirez: API pricing is Monopoly money — the only real metric is joules, and we can't see them — mitsuhiko · 2026-09-19
- Domestic micro-datacentres strapped to water tanks cut UK heating bills by £10-15/month — nordicinst · 2026-09-19