Open-source project runs 120B MoE models on smartphones at 6 tokens/s
dai_app · reddit · 2026-07-30
Developer daiapp created bigedgeonmoe, an open-source codebase that runs massive MoE models (Qwen 35B to 120B) on mobile devices or consumer PCs. Qwen 35B (Q4) achieves 6 tokens/s on a mid-range phone with 12GB RAM. It's modular on top of llama.cpp, supporting any model/quantization, and new architectures can be registered with one line of code.
More from Infra
- Analyst: AI Fundamentals Unchanged Amidst Severe Compute Deficit and Low Penetration — BenBajarin · 2026-07-31
- AI Buildout Validates Kleiner Perkins' Cleantech Fund 20 Years Later — matt_slotnick · 2026-07-31
- China's DUV Lithography: Not an ASML Killer, But an Iteration Loop — demian_ai · 2026-07-31
- AWS Guide: Building Inference Meta-Monitoring with SageMaker AI and Quick — AWS ML Blog · 2026-07-31
- Help needed: Native 64k+ context GGUF model for llama.cpp — Inner-End7733 · 2026-07-31
- OpenAI GPT-5.6 Models Hit Amazon Bedrock with Explicit Prompt Caching — AWS ML Blog · 2026-07-31