27B Model Compressed to Run on iPhone

xiaohu · x · 2026-07-15

PrismML applied model compression to Qwen3.6-27B, shrinking the original 54GB model down to 3.9–5.9GB to enable local execution on smartphones.

They compared the compressed versions against the original across 15 tests under an "intense thinking mode":

Regarding performance, the author claims speeds up to 163 token/s on an RTX 5090 and 87 token/s on an Apple M5 Max.

Related event: Bonsai 27B: The first 27B model that runs on phones(15 posts)→

Original post →

More from Infra

Infra channel →