Bonsai 2 squeezes Qwen 27B into 6GB with ternary weights, but fails agentic tasks
Prompt Engineering · youtube · 2026-09-28
- PrismML's Bonsai 2 retrains every weight to ternary values (−1/0/+1), fitting Qwen3.8 27B into a 6GB file while claiming 98% of full-model performance.
- The reviewer found it matched the full model on short chat/coding tasks, but in an agentic loop building an ISS tracker on a 3D globe, the full model shipped a working app while Bonsai ran the same search 114 times without writing a single line.
- Takeaway: compressed ternary models are fine for laptop chat and quick coding help, but test them on your own long-running agent tasks before trusting the benchmarks.
Related event: Bonsai 2 Compresses Qwen 27B to 6GB but Falters on Long Tasks(2 posts)→
More from Infra
- Exploit Summit Montreal recap: Gamma tokens, iota SDK, $12M run rate for Targon — markjeffrey · 2026-09-29
- Bain says AI must earn $6T a year by 2031 — matching all global IT spending today — sanjaykalra · 2026-09-29
- On DGX Spark, bf16 beats int8 convrot: H3 video gen 272s vs 287s in real tests — dtdisapointingresult · 2026-09-29
- BAAI's MALA attention allocates its own compute, cutting 128K training latency 2.2x — BAAI · 2026-09-29
- BAAI's CoWA attention cuts training latency 7.4x while matching FullAttn quality to 32B — BAAI · 2026-09-29
- Databricks Tops All 4 NVIDIA SOL-ExecBench Kernel Tracks Using AI Agents for ~$70K — Yuchenj_UW · 2026-09-29