Streaming 120B MoE Inference on Android

dai_app · reddit · 2026-07-18

Using a OnePlus 15R Android phone with roughly 11GB of available memory, the author ran several large models—including GPT-OSS-120B, Qwen 30B, and Gemma 26B—entirely on CPU and flash storage, achieving speeds of 1–5 tok/s.

Core Approach

Engineering Details

Benchmarks

Challenges

The real difficulty wasn't streaming reads, but Android reclaiming resident weights under memory pressure, causing continuous re-paging during generation. The author notes this consumed the bulk of the development effort. The project is Apache-2.0 licensed and offers a pre-compiled APK.

Related event: Android Demo Runs 120B-Class MoE(2 posts)→

Original post →

More from Infra

Infra channel →