Andrew Chen: strong LLMs are far from running on phones, on-device AI faces bandwidth, heat and model-size hurdles

andrewchen · x · 2026-09-21

Andrew Chen argues that strong LLMs remain far from native on-device execution: slow memory bandwidth, only highly quantized MoE models fit, and power/heat remain issues. Next-gen mobile NPUs target only modestly sized LLMs.

He sees major mobile UX opportunities — notifications, typing, inboxes, calendars — that would benefit from fast, cheap AI decision models, and floats the idea of baking an older-but-useful model directly into phone hardware, wondering if Jevons-like dynamics could accelerate the shift.

Original post →

More from Embodied

Embodied channel →