PrismML's Bonsai 2 shrinks Qwen3.8 27B to 5.9GB, keeps 98.2% capability, runs on a 5090
dl_weekly · x · 2026-09-28
PrismML launched Bonsai 2 27B, an Apache-2.0 model that compresses Qwen3.8 27B from 56GB (16-bit) to 5.9GB using ternary weights (+1/0/-1), retaining 98.2% capability and running at 143 tokens/s on a consumer GeForce 5090. Benchmarks stay close to the original: agentic/tool-calling 77.6 vs 79.8, coding 81.6 vs 82.2, knowledge & reasoning 82.7 vs 81.3 — showing ternary compression can bring large models onto consumer PCs and high-end mobile devices.
More from Infra
- Fireworks' Ember-1 post-trains Kimi K3 to reason 40% more concisely at same quality — isidentical · 2026-09-28
- One AMD driver flag boosts dual-GPU Vulkan LLM inference up to 4x — tabletuser_blogspot · 2026-09-28
- Developer Plans to Let Codex Pick Which Tests Run, Slashing CI Costs — sull · 2026-09-28
- Free online guide covers LLMs from first principles to local deployment — JFPuget · 2026-09-28
- Is a vector database enough for production AI agents? Reddit debates storage design — OkShirt9372 · 2026-09-28
- Pre-training a Foundation Model Whose Tokenizer Is PTX, Not Natural Language — rickasaurus · 2026-09-28