Open-source project runs 120B MoE models on smartphones at 6 tokens/s

dai_app · reddit · 2026-07-30

Developer daiapp created bigedgeonmoe, an open-source codebase that runs massive MoE models (Qwen 35B to 120B) on mobile devices or consumer PCs. Qwen 35B (Q4) achieves 6 tokens/s on a mid-range phone with 12GB RAM. It's modular on top of llama.cpp, supporting any model/quantization, and new architectures can be registered with one line of code.

Original post →

More from Infra

Infra channel →