FULL STORY
Bonsai 2 Ternary Quantization: From Viral Launch to Real-World Struggles
PrismML's Bonsai 2 launched with a 6GB ternary-quantized 27B model claiming 98.2% performance retention, but real-world tests ten days later showed it failing on longer tasks, sparking debate over ternary quantization's reliability.
2026-09-18 ~ 2026-09-28 · 2 episodes · 22 posts
Episode 1 · PrismML's Ternary Bonsai 2 27B Shrinks to 5.9GB, Retains 98.2% Performance (2026-09-18, 20 posts)
On September 18, PrismML released Ternary Bonsai 2 (27B) on Hugging Face, which quickly hit the trending charts and launched on OpenRouter the next day. Built on Qwen3.8-27B's hybrid-attention architecture with no architectural changes, the model ternarizes weights to {-1, 0, +1} with FP16 group scaling (1.76 bits per effective weight), shrinking to 5.9GB (model card: 5.95GB) versus 54GB FP16—about 9x smaller—while claiming 98.2% of aggregate benchmark performance and running on consumer hardware and even WebGPU browsers.
Confirmed
- Ternary quantization on Qwen3.8-27B with prismhadamardqwen35 config, 1.76 bits/weight effective, 5.9GB total vs 54GB FP16
- Claimed 98.2% of FP16 aggregate benchmark scores
- Multiple formats: GGUF (llama.cpp-ready), MLX 2-bit, ONNX; runs on M5 Max, RTX 3080, WebGPU browsers, and reportedly phones
- @MaziyarPanahi ran it locally on Mac, reading medication-label changes in 5.7 seconds; @airesearch12 noted a community CUDA patch yielding a further 32% speedup
- Apache 2.0 licensed, agentic-focused; launched on OpenRouter Sep 19 with 262K context, coding/math/tool-use/vision support, thinking mode on by default (8.5GB weights on OpenRouter)
- Second-generation Bonsai, two months after the first, same footprint with improved post-quantization quality
Unconfirmed
- The 98.2% retention figure is vendor-reported and not independently verified, as flagged by @airesearch12 and @MaziyarPanahi
Why it matters
Ternary quantization could drastically cut memory and inference costs, putting a 27B model on consumer devices, phones, and browsers. If quality loss proves minimal upon independent verification, this could reshape how large models are distributed and used; the one-day hop to OpenRouter also shows fast deployment from open weights to hosted service.
- Ternary Bonsai 2 27B: 9x smaller, 98.2% of full-precision performance, Apache 2.0 — cephaloform · 2026-09-18
- Ternary Bonsai 2 27B: 9x smaller than Qwen3.8 27B at 98.2% performance, Apache 2.0 — cephaloform · 2026-09-18
- Ternary Bonsai 2: 27B Model Under 6GB Runs In-Browser on WebGPU, Keeps 98.2% Quality — xenovatech · 2026-09-18
- Ternary Bonsai 2 hits Hugging Face: 27B ternary reasoning model under 6GB, runs in browser — cephaloform · 2026-09-18
- Bonsai 2 27B ships with ternary weights: 5.95GB model hits 98.2% of FP16 benchmarks — airesearch12 · 2026-09-18
- Ternary-Bonsai-2-27B, a 2-bit ternary model for on-device inference, trends on Hugging Face — prism-ml · 2026-09-18
- Ternary Bonsai 2 27B: 9x smaller than Qwen3.8 27B, keeps 98.2% of performance — TheZachMueller · 2026-09-18
- Ternary Bonsai 2 27B released: 9x smaller, 98.2% of full precision, just 5.9 GB under Apache 2.0 — cephaloform · 2026-09-18
- Ternary Bonsai 2 27B: 5.9GB open model keeps 98.2% of full-precision performance — gajesh · 2026-09-18
- PrismML releases ternary Bonsai 2 27B: 9x smaller, keeps 98.2% of Qwen3.8 27B performance — pcuenq · 2026-09-18
- Ternary-Bonsai-2-27B 2-bit MLX quantized model trends on Hugging Face — prism-ml · 2026-09-18
- PrismML launches ternary Bonsai 2 27B: 5.9GB, 9x smaller, retains 98.2% of full-precision performance — airesearch12 · 2026-09-18
- Bonsai 2 27B: 9x smaller ternary model keeps 98.2% of benchmarks, runs on a Mac — MaziyarPanahi · 2026-09-18
- Ternary Bonsai-2 27B reads medication lists locally on Mac in 5.7 seconds — MaziyarPanahi · 2026-09-18
- Ternary Bonsai 2 27B: a 5.9GB open model keeping 98.2% of full-precision performance — ivan_bezdomny · 2026-09-18
- Ternary Bonsai 2 27B is 9x smaller than Qwen3.8 27B while keeping 98.2% of its benchmark performance — alexcovo_eth · 2026-09-19
- PrismML's Ternary Bonsai 2 27B is 9x smaller at 5.9GB while keeping 98.2% benchmark performance — alexcovo_eth · 2026-09-19
- PrismML Launches Bonsai 2 27B: 9x Smaller Than Qwen3.8, Keeps 98.2% Performance — ivan_bezdomny · 2026-09-19
- PrismML's ternary-compressed Bonsai 2 27B lands on OpenRouter at just ~8.5GB — gajesh · 2026-09-19
- 1.58-bit ternary 27B model runs on 12GB cards; CUDA patch adds 28-32% speed — airesearch12 · 2026-09-19
Episode 2 · Bonsai 2 Compresses Qwen 27B to 6GB but Falters on Long Tasks (2026-09-28, 2 posts)
PrismML's Bonsai 2 ternary compression shrinks Qwen 27B to a 6GB file while claiming 98% performance, but real-world testing shows it matches the full model only on short tasks and collapses on longer agent workflows.
- Bonsai 2 squeezes Qwen 27B into 6GB with ternary weights, but fails agentic tasks — Prompt Engineering · 2026-09-28
- Bonsai 2 tested: 98% of Qwen 27B in 6GB, but zero working agent builds — Prompt Engineering · 2026-09-28