Qualcomm's Next Hexagon NPU: 50% More Shared Memory, 30B MoE Models on a Phone

ryanshrout · x · 2026-09-11

Third pre-Summit deep dive: Qualcomm's next Hexagon NPU adds an Element Accelerator combining vector and scalar extensions for transformer workloads, a shared memory pool 50% larger than Snapdragon 8 Elite Gen 5, and support for MoE models up to 30B parameters.

Claimed gains vs 8 Elite Gen 5: up to 50% higher INT4 pre-fill performance, time-to-first-token as low as 1.5s.

The article argues MoE on a phone is fundamentally a storage problem: a 30B MoE keeps tens of billions of parameters available while activating only 3B routed parameters per token, keeping compute within a phone power budget while behaving like a much larger model.

Original post →

More from Embodied

Embodied channel →