Qwen 3.6 35B MoE runs on a Xiaomi 12 Pro with 12GB RAM at 2.4 tok/s

Aromatic_Ad_7557 · reddit · 2026-07-24

A Reddit user reports running Qwen 3.6 35B MoE in Q4KM on a modified Xiaomi 12 Pro with just 12GB of RAM.

Using BigMoeOnEdge, the setup reaches 2.428 tok/s, with TTFT around 19.9s, 8k context, and heavy CPU usage plus paging. The post includes detailed token/compute/cache stats and shows that a very large MoE model can run locally on a phone-class edge device, albeit with clear memory and latency trade-offs.

Original post →

More from Infra

Infra channel →