Running a 27B Model on an iPhone

Prompt Engineering · youtube · 2026-07-15

The video breaks down how Prism ML compressed a 27B parameter Qwen-series model to run on an iPhone 17 Pro at roughly 11 tokens/s:

Related event: Quantized Qwen 27B Model Runs on iPhone(2 posts)→

Original post →

More from Infra

Infra channel →