Qwen 3.6 35B MoE runs on a Xiaomi 12 Pro with 12GB RAM at 2.4 tok/s
Aromatic_Ad_7557 · reddit · 2026-07-24
A Reddit user reports running Qwen 3.6 35B MoE in Q4KM on a modified Xiaomi 12 Pro with just 12GB of RAM.
Using BigMoeOnEdge, the setup reaches 2.428 tok/s, with TTFT around 19.9s, 8k context, and heavy CPU usage plus paging. The post includes detailed token/compute/cache stats and shows that a very large MoE model can run locally on a phone-class edge device, albeit with clear memory and latency trade-offs.
More from Infra
- Pydantic AI says one agent can spawn 40 test suites, and the laptop fans prove it — AAAzzam · 2026-07-24
- Fluidstack raises $830M Series A at a $7.5B valuation — MxMnr · 2026-07-24
- Google’s capex guide now exceeds all spending from founding through 2021 — ivan_bezdomny · 2026-07-24
- Intel says advanced packaging backlog is building, with billions in sight — BenBajarin · 2026-07-24
- Intel says Q2 data center sales rose 59% as demand outran supply — Polymarket · 2026-07-24
- AI recommendations may steer shoppers toward luxury picks before cheaper options — PierceLilholt · 2026-07-24