Bonsai 2 tested: 98% of Qwen 27B in 6GB, but zero working agent builds
Prompt Engineering · youtube · 2026-09-28
The Prompt Engineering YouTube channel tested PrismML's Bonsai 2, a ternary compression scheme claiming to keep 98% of the Qwen 27B model's performance in a 6GB file.
Findings:
- On short tasks (thinking off, same agent harness) it's a tie: 48.6 vs 47.1
- On six long agentic builds, full Qwen scored 35/60 while Bonsai 2 produced zero working apps, once repeating the same search 114 times
- The video explains how ternary compression works, what PrismML's own whitepaper says about agentic benchmarks, and what trace signals to watch before trusting a compressed model with agent work
Takeaway: compressed models may be fine for chat, but think twice before handing them long-horizon agent tasks.
Related event: Bonsai 2 Compresses Qwen 27B to 6GB but Falters on Long Tasks(2 posts)→
More from Models
- Early Sonnet 5.5 impressions: ~5x faster bug fixes at high effort — burkov · 2026-09-29
- Why 4o feels different: thread argues native omni training, not capability, shapes model personality — RileyRalmuto · 2026-09-29
- Sonnet 5.5 effort settings make no difference in 15-task coding test: 9/15 at low, medium and high — every · 2026-09-29
- PrunaAI claims its text-to-video modes sit on DesignArena Pareto frontiers — guennemann · 2026-09-29
- Reddit user: Opus 5.5 silently falls back to Opus 5 on nearly every prompt — fishcat_catfish · 2026-09-29
- Sonnet 5.5 clones open-source editor Proof at low effort, joining elite group of just four models — every · 2026-09-29