PrismML Reiterates Mobile Compression of Qwen27B

xiaohu · x · 2026-07-15

PrismML continues to highlight its success in compressing Qwen3.6-27B for mobile execution: the roughly 54GB model has been reduced to 3.9–5.9GB, claiming to retain most of the original capabilities in testing.

The post reiterates performance metrics: the Ternary version retains about 95% of its power, while the 1-bit mobile version retains roughly 90%. It also mentions inference speeds on the RTX 5090 and Apple M5 Max.

Related event: Bonsai 27B: The first 27B model that runs on phones(15 posts)→

Original post →

More from Infra

Infra channel →