Quantized Qwen 27B Model Runs on iPhone
Prism ML's Bonsai 27B project uses extreme quantization to compress a Qwen-based 27B model to 3.9GB. This allows it to run locally on an iPhone 17 Pro at about 11 tokens/s and load easily on a single RTX 3090.
2026-07-15 ~ 2026-07-17 · 2 related posts
- Running a 27B Model on an iPhone — Prompt Engineering · 2026-07-15
- Qwen 27B Compressed Down to 3.9GB — max_paperclips · 2026-07-17