Ternary Bonsai 2: 27B Model Under 6GB Runs In-Browser on WebGPU, Keeps 98.2% Quality
xenovatech · reddit · 2026-09-18
Ternary Bonsai 2 (27B) is now on Hugging Face. Derived from the Qwen3.8-27B hybrid-attention architecture (unchanged), it uses ternary weights to shrink the model to under 6GB—9x smaller than FP16 while retaining 98.2% of the intelligence, per the model card. Its small footprint lets it run locally in-browser via WebGPU, with a live demo and collection available.
Related event: Ternary Bonsai 2 27B Released: 9x Smaller, Keeps 98.2% Performance(5 posts)→
More from Infra
- Brad Gerstner at All-In Summit: who pays for AI CapEx, the gigawatt gap and semis eating the Nasdaq — DavidSacks · 2026-09-18
- Third-party audit reproduces Gensyn open-1b training step bit-for-bit — benfielding · 2026-09-18
- Anthropic Open-Sources Claude-Written GPU Optimizations Speeding 30+ Biomolecular Models ~4x — ResultBackground2450 · 2026-09-18
- Spotify: 777M users, 11-12M requests/sec — how AI changed its quality playbook — rseroter · 2026-09-18
- 605 new Linux kernel CVEs disclosed in one day, on top of 276 the day before — jedisct1 · 2026-09-18
- DeepSeek V4.1 Flash hits 532 tokens/s on Inco, fastest output on Artificial Analysis — songhan_mit · 2026-09-18