Bonsai 2 27B ships with ternary weights: 5.95GB model hits 98.2% of FP16 benchmarks
airesearch12 · x · 2026-09-18
Bonsai 2 27B is out with ternary weights: the language model is just 5.95GB versus 54GB in FP16, while scoring 98.2% of the FP16 Qwen3.8 27B benchmarks.
Key details:
- Runs at 55 tok/s via WebGPU on an M5 Max, with a browser demo available
- 27B reasoning model with vision, tool use, and agentic support
- Open weights released
More from Infra
- Cactus Releases Needle 3: An 8-29MB Foundation Model Running 4k tokens/s on a Raspberry Pi 5 — airesearch12 · 2026-09-18
- Bittensor Subnet 105 Beam launches as a low-cost high-speed bandwidth network for AI — markjeffrey · 2026-09-18
- DeepSeek v4.1 flash TTFT comparison: Together AI crushes rivals on pre-warmed queries — zhyncs42 · 2026-09-18
- Huawei's 100K-card Super Cluster can train a 10T-param model on 100T tokens in 30 days — teortaxesTex · 2026-09-18
- MLX Community Makes Qwen 3.8 Flash Nearly 2x Faster on Apple Silicon, License Blocks Launch — gajesh · 2026-09-18
- CoreWeave Brings Multi-Rack NVIDIA Vera Rubin NVL72 Cluster Online — Beth_Kindig · 2026-09-18