Run a 92GB Model on a 12GB Phone: BigMoeOnEdge Breaks Edge MoE Limits
dai_app · reddit · 2026-08-05
A developer demonstrated BigMoeOnEdge, an open-source engine that successfully ran the massive 92GB DeepSeek V4 Flash IQ2M model on a mid-range Android phone with 12GB of RAM at 1 token/s.
- Technical Breakthrough: Leveraging the modularity of llama.cpp, the engine proves that even exceptionally large MoE models can achieve responsiveness on resource-constrained mobile devices or consumer PCs.
- Compatibility: Beyond DeepSeek, the engine supports various models including Gemma 26B and Qwen 3, with some users reportedly running 397B models on their phones via this framework.
- While 1 token/s is not yet practical for daily use, it serves as a solid proof-of-concept for ultra-large edge deployment, requiring just one line of code to run any supported large MoE model.
More from Infra
- OpenAI Engineer Details WebRTC Improvements Powering GPT-Live — juberti · 2026-08-05
- Musk: Starmind Satellite Design Will Improve Ground Data Centers — elonmusk · 2026-08-05
- SpaceX and NVIDIA Plan Up to One Million AI Compute Satellites — ns123abc · 2026-08-05
- 20B Model Maple-Preview Runs at 200+ tokens/s on Mac Mini, Solves IMO Math — tylerbruno05 · 2026-08-05
- Open Source AI Hosting Hits Compute Wall: DeepSeek Flash Offline Amid GPU Crunch — bindureddy · 2026-08-05
- Full Kimi K3 model runs on 16x GB10 cluster at 20+ TPS — ciprianveg · 2026-08-05