Bonsai 27B: The first 27B model that runs on phones
On July 15, PrismML released Bonsai 27B, claiming it as the first 27B-class model runnable on phones. Built on Alibaba's Qwen3.6 27B and compressed via extreme low-bit quantization from roughly 54GB to a few GB, it pushes 27B-class capability — once confined to the cloud or high-end workstations — onto phones, browsers and low-VRAM devices, drawing quick attention from the local-inference community.
Key details
@thoquz notes that PrismML shipped both Hugging Face weights and a GGUF version for easy local loading, while @xenovatech reports the model uses a 1-bit dense LLM architecture with custom WebGPU kernels to run directly in the browser. On supported platforms, @awnihannun (forward) lists iPhone 17 Pro, iPhone Air and some iPads; @tcarambat and @immersive-matthew emphasize fitting a 27B-class model into about 10GB of RAM or 12GB of VRAM.
Performance and claims
Reported compressed sizes range from 3.8GB to 3.9–5.9GB across posts. @xiaohu and @airesearch12 relay PrismML's testing: across 15 comparisons the Ternary (1.58-bit) variant retains about 95% of capability and the 1-bit variant about 90%. @pbaylies (forward) adds Ternary-Bonsai-27B llama-bench throughput figures on a single 3090 at 32k/64k/128k context (about 956 at 32k). Multiple forwards also mention support for multi-step reasoning, structured tool calls and long-context workflows; @victormustar (forward) adds that evaluation used Fable 5 and GPT 5.6 Sol.
Timeline and impact
@pmttyji reported on July 13 that PrismML had compressed Qwen 3.6-27B to run on iPhone 17 Pro, with download opening the following week; by July 15 multiple posts confirm Bonsai 27B's official release. If the size and capability-retention figures hold, Bonsai 27B represents an aggressive localization route: using extreme low-bit compression to bring 27B-class models into phone and browser scenarios.
2026-07-13 ~ 2026-07-15 · 15 related posts
- [source] Qwen 27B Compressed to Run on iPhone — pmttyji · 2026-07-13
- 27B Ternary Model Runs Within 10GB Memory — tcarambat · 2026-07-15
- [source] Bonsai 27B Released and Available for Local Execution — thoquz · 2026-07-15
- 1-Bit LLM Runs Locally in the Browser — xenovatech · 2026-07-15
- 27B Model Compressed to 3.8GB — victormustar · 2026-07-15
- Qwen3.6 27B Retains Most Intelligence After 1.58-bit Quantization — airesearch12 · 2026-07-15
- 27B Multimodal Model Now Runs on Mobile Phones — _arohan_ · 2026-07-15
- 27B Model Can Now Run on Mobile Phones — yogthos · 2026-07-15
- 27B Model Can Now Run Directly on Mobile Phones — awnihannun · 2026-07-15
- [source] 27B Model Compressed to Run on iPhone — xiaohu · 2026-07-15
- PrismML Reiterates Mobile Compression of Qwen27B — xiaohu · 2026-07-15
- Qwen3.6 27B Fits Into 12GB VRAM — immersive-matthew · 2026-07-15
- Ternary-Bonsai-27B Local Inference Tested — pbaylies · 2026-07-15
2 near-duplicate retellings: dair_ai · zacharynado