Qualcomm's next-gen Hexagon NPU runs 30B MoE models with 32K context on-device

lee_stott · x · 2026-09-11

Qualcomm detailed its next-gen Hexagon NPU for Snapdragon devices: a new Element Accelerator for transformers, +50% shared memory, 32K context window, the ability to run 30B-parameter MoE models (3B active per token), and +50% INT4 prefill speed — a big step for agentic on-device AI.

Related event: Qualcomm's next Hexagon NPU to run 30B MoE models on-device(2 posts)→

Original post →

More from Embodied

Embodied channel →