PrismML Launches Bonsai 2 27B: 9x Smaller Than Qwen3.8, Keeps 98.2% Performance
ivan_bezdomny · x · 2026-09-19
PrismML open-sourced Ternary Bonsai 2 27B, a ternary-weight compression of Qwen3.8 27B that shrinks memory to 5.9GB while retaining 98.2% of aggregate benchmark performance, under Apache 2.0.
- Runs fully locally in-browser via WebGPU or on consumer GPUs like a 12GB RTX 3060 (220K context, 35 tokens/s decode measured)
- Targets agentic workloads: multimodal reasoning, agentic coding, and long-horizon tool use, with fewer silent failures and derailments on multi-step tasks vs the first 27B release
- Team demos local computer-use workflows on an RTX 5090
Related event: PrismML's Ternary Bonsai 2 27B Shrinks to 5.9GB, Retains 98.2% Performance(20 posts)→
More from Infra
- SemiAnalysis: Why GLM-5.3 Sparse Attention Doesn't Cut HBM Memory Capacity Needs — burny_tech · 2026-09-29
- Exploit Summit Montreal recap: Gamma tokens, iota SDK, $12M run rate for Targon — markjeffrey · 2026-09-29
- Bain says AI must earn $6T a year by 2031 — matching all global IT spending today — sanjaykalra · 2026-09-29
- On DGX Spark, bf16 beats int8 convrot: H3 video gen 272s vs 287s in real tests — dtdisapointingresult · 2026-09-29
- BAAI's CoWA attention cuts training latency 7.4x while matching FullAttn quality to 32B — BAAI · 2026-09-29
- BAAI's MALA attention allocates its own compute, cutting 128K training latency 2.2x — BAAI · 2026-09-29