Ternary Bonsai 2: 27B Model Under 6GB Runs In-Browser on WebGPU, Keeps 98.2% Quality

xenovatech · reddit · 2026-09-18

Ternary Bonsai 2 (27B) is now on Hugging Face. Derived from the Qwen3.8-27B hybrid-attention architecture (unchanged), it uses ternary weights to shrink the model to under 6GB—9x smaller than FP16 while retaining 98.2% of the intelligence, per the model card. Its small footprint lets it run locally in-browser via WebGPU, with a live demo and collection available.

Related event: Ternary Bonsai 2 27B Released: 9x Smaller, Keeps 98.2% Performance(5 posts)→

Original post →

More from Infra

Infra channel →