FULL STORY

Bonsai 2 Ternary Quantization: From Viral Launch to Real-World Struggles

PrismML's Bonsai 2 launched with a 6GB ternary-quantized 27B model claiming 98.2% performance retention, but real-world tests ten days later showed it failing on longer tasks, sparking debate over ternary quantization's reliability.

2026-09-18 ~ 2026-09-28 · 2 episodes · 22 posts

Episode 1 · PrismML's Ternary Bonsai 2 27B Shrinks to 5.9GB, Retains 98.2% Performance (2026-09-18, 20 posts)

On September 18, PrismML released Ternary Bonsai 2 (27B) on Hugging Face, which quickly hit the trending charts and launched on OpenRouter the next day. Built on Qwen3.8-27B's hybrid-attention architecture with no architectural changes, the model ternarizes weights to {-1, 0, +1} with FP16 group scaling (1.76 bits per effective weight), shrinking to 5.9GB (model card: 5.95GB) versus 54GB FP16—about 9x smaller—while claiming 98.2% of aggregate benchmark performance and running on consumer hardware and even WebGPU browsers.

Confirmed

  • Ternary quantization on Qwen3.8-27B with prismhadamardqwen35 config, 1.76 bits/weight effective, 5.9GB total vs 54GB FP16
  • Claimed 98.2% of FP16 aggregate benchmark scores
  • Multiple formats: GGUF (llama.cpp-ready), MLX 2-bit, ONNX; runs on M5 Max, RTX 3080, WebGPU browsers, and reportedly phones
  • @MaziyarPanahi ran it locally on Mac, reading medication-label changes in 5.7 seconds; @airesearch12 noted a community CUDA patch yielding a further 32% speedup
  • Apache 2.0 licensed, agentic-focused; launched on OpenRouter Sep 19 with 262K context, coding/math/tool-use/vision support, thinking mode on by default (8.5GB weights on OpenRouter)
  • Second-generation Bonsai, two months after the first, same footprint with improved post-quantization quality

Unconfirmed

  • The 98.2% retention figure is vendor-reported and not independently verified, as flagged by @airesearch12 and @MaziyarPanahi

Why it matters

Ternary quantization could drastically cut memory and inference costs, putting a 27B model on consumer devices, phones, and browsers. If quality loss proves minimal upon independent verification, this could reshape how large models are distributed and used; the one-day hop to OpenRouter also shows fast deployment from open weights to hosted service.

Episode 2 · Bonsai 2 Compresses Qwen 27B to 6GB but Falters on Long Tasks (2026-09-28, 2 posts)

PrismML's Bonsai 2 ternary compression shrinks Qwen 27B to a 6GB file while claiming 98% performance, but real-world testing shows it matches the full model only on short tasks and collapses on longer agent workflows.