Bonsai 2 challenge: Qwen 27B compressed 10x already 140% faster on Mac, contest open
gajesh · x · 2026-09-26
The MLX.fast community launched a one-week Ternary Bonsai 2 challenge to optimize PrismML's Bonsai 2 — a near-lossless compression of Qwen 3.8 27b to roughly 10% of the original size, fitting on 16GB Apple devices and some phones.
- The community doubled Bonsai's speed within 12 hours; plenty of headroom remains
- Current record: 140.7% speedup over baseline on an M5 Mac (158.2 decode tok/s, 953.5 prefill tok/s) across 17 promoted submissions from 7 solvers
- Score = prefill^0.25 · decode^0.75, so decoding carries 75% of the weight
- Contest runs until October 1 and is open to everyone
More from Infra
- Germany and the Netherlands put €40m into a challenge to design AI chips with AI — VraserX · 2026-09-26
- AMD Claims One Factory Box Can Run 2.3x More Software Workers — shashib · 2026-09-26
- Token Prices Fall via Subsidies, Moore's Law via Manufacturing — Different Drivers — HanchungLee · 2026-09-26
- KV cache transplants on Qwen3.8-27B: start at Q6, hand off to Q3, beat static quant — wadeAlexC · 2026-09-26
- Google open-sources GKE agentic migration for AI-assisted EKS-to-GKE moves — rseroter · 2026-09-26
- Musk: China's electricity output now exceeds the US, Europe and India combined — XFreeze · 2026-09-26