Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning
cephaloform · x · 2026-09-23
Ternary Bonsai 2 27B is out: ternary-weight compression shrinks a 27B model to just 5.9GB of weights, versus 54GB for full-precision Qwen3 27B.
The author benchmarked how much reasoning survived compression against full-precision Qwen3 27B and Gemma 4 12B QAT (7GB):
- Setup: no internet, no tools, minimal harness for coding problems; default llama.cpp serving configs, extra-high reasoning for Bonsai/Qwen
- On IMO 2026 (released after all training cutoffs), each model got a 131k thinking context before being forced to answer
- Result: both Qwen and Bonsai land in the upper end of the human bronze range (16-22), with Bonsai retaining 95% of base-model performance
The writeup discusses where compression preserves performance and where gaps remain.
More from Infra
- Inside the DJI teardown: Sarah Guo says Shenzhen's manufacturing knowledge can't be scraped into a model — bookwormengr · 2026-09-23
- Your data stack is about to get less forgiving: agents need data that's true now — bigdata · 2026-09-23
- Qwen 3.6 35B-A3B Q6 hits ~50 tok/s on a 128GB Strix Halo — what's the best local model now? — jankeydankey · 2026-09-23
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Qwen 27B runs 24hr unattended on one RTX5090, builds full Postgres-SpringBoot-React spreadsheet app — anglepoiselife · 2026-09-23