8B Model Hits 60 Tokens/sec on Phone CPU Alone, No GPU or NPU Needed
const_reborn · x · 2026-09-29
At the Exploit conference, @jondurbin reported running an 8B model entirely on a phone's CPU — no GPU or NPU — at roughly 60 tokens per second.
The test used real devices rented through Qualcomm Device Cloud, and notably was measured before any optimization work. That suggests meaningful headroom for on-device LLM performance once mobile-specific quantization and scheduling tweaks land.
More from Infra
- The endless model-hopper cycle: Claude's success overloads servers, users drift back to GPT — haider1 · 2026-09-29
- AI detector startup Pangram saved 2-3 engineer-years and cut compute costs 50-60% — LightningAI · 2026-09-29
- Batam Data Center Strains Residents' Water Supply While Its Desalination Plant Remains on Paper — AryHHAry · 2026-09-29
- Nvidia and AMD lobby Trump to keep China chip sales flowing, opposing AI OVERWATCH Act — pstAsiatech · 2026-09-29
- Chinese Phone Makers Raise Flagship Prices by Up to $446 as Memory Costs Surge — pstAsiatech · 2026-09-29
- Running 1000s of enterprise RAG indexes: a year of lessons in incremental sync — srnsnemil · 2026-09-29