PrismML releases ternary Bonsai 2 27B: 9x smaller, keeps 98.2% of Qwen3.8 27B performance
pcuenq · x · 2026-09-18
PrismML announced Ternary Bonsai 2 27B, a ternary-quantized model based on Qwen3.8 27B that is 9x smaller than its full-precision counterpart (about 5.9 GB) while retaining 98.2% of the parent model's aggregate benchmark performance.
- Two months after the first Bonsai 27B, the biggest change is quality: the gap to full precision narrowed materially, with strong gains in agentic coding, multimodal reasoning, and long-horizon tool use
- Runs at 55 tok/s on an M5 Max at under 6 GB
- Released today under Apache 2.0
Related event: PrismML's Ternary Bonsai 2 27B Shrinks to 5.9GB, Retains 98.2% Performance(20 posts)→
More from Infra
- Brookings paper projects $10.3T in AI infrastructure investment by 2032, flags hidden financing risks — VraserX · 2026-09-29
- SemiAnalysis: Why GLM-5.3 Sparse Attention Doesn't Cut HBM Memory Capacity Needs — burny_tech · 2026-09-29
- Exploit Summit Montreal recap: Gamma tokens, iota SDK, $12M run rate for Targon — markjeffrey · 2026-09-29
- Bain says AI must earn $6T a year by 2031 — matching all global IT spending today — sanjaykalra · 2026-09-29
- On DGX Spark, bf16 beats int8 convrot: H3 video gen 272s vs 287s in real tests — dtdisapointingresult · 2026-09-29
- BAAI's CoWA attention cuts training latency 7.4x while matching FullAttn quality to 32B — BAAI · 2026-09-29