PrismML's ternary-compressed Bonsai 2 27B lands on OpenRouter at just ~8.5GB
gajesh · x · 2026-09-19
PrismML's reasoning model Ternary Bonsai 2 27B is now live on OpenRouter. Derived from Qwen3.8-27B, it supports coding, math, tool calling, and image understanding with a 262K context window, thinking by default at xhigh reasoning effort.
Its key feature is ternary compression: language-model weights shrink to roughly 8.5GB while retaining 98.2% of the base model's average score across PrismML's 14 thinking-mode benchmarks, enabling efficient inference on consumer hardware. Pricing is $0.075/$0.50 per 1M input/output tokens.
Related event: PrismML's Ternary Bonsai 2 27B Shrinks to 5.9GB, Retains 98.2% Performance(20 posts)→
More from Infra
- Brookings paper projects $10.3T in AI infrastructure investment by 2032, flags hidden financing risks — VraserX · 2026-09-29
- SemiAnalysis: Why GLM-5.3 Sparse Attention Doesn't Cut HBM Memory Capacity Needs — burny_tech · 2026-09-29
- Exploit Summit Montreal recap: Gamma tokens, iota SDK, $12M run rate for Targon — markjeffrey · 2026-09-29
- Bain says AI must earn $6T a year by 2031 — matching all global IT spending today — sanjaykalra · 2026-09-29
- On DGX Spark, bf16 beats int8 convrot: H3 video gen 272s vs 287s in real tests — dtdisapointingresult · 2026-09-29
- BAAI's CoWA attention cuts training latency 7.4x while matching FullAttn quality to 32B — BAAI · 2026-09-29