8B Model Hits 60 Tokens/sec on Phone CPU Alone, No GPU or NPU Needed

const_reborn · x · 2026-09-29

At the Exploit conference, @jondurbin reported running an 8B model entirely on a phone's CPU — no GPU or NPU — at roughly 60 tokens per second.

The test used real devices rented through Qualcomm Device Cloud, and notably was measured before any optimization work. That suggests meaningful headroom for on-device LLM performance once mobile-specific quantization and scheduling tweaks land.

Original post →

More from Infra

Infra channel →