Qualcomm's Next Hexagon NPU: 50% More Shared Memory, 30B MoE Models on a Phone
ryanshrout · x · 2026-09-11
Third pre-Summit deep dive: Qualcomm's next Hexagon NPU adds an Element Accelerator combining vector and scalar extensions for transformer workloads, a shared memory pool 50% larger than Snapdragon 8 Elite Gen 5, and support for MoE models up to 30B parameters.
Claimed gains vs 8 Elite Gen 5: up to 50% higher INT4 pre-fill performance, time-to-first-token as low as 1.5s.
The article argues MoE on a phone is fundamentally a storage problem: a 30B MoE keeps tens of billions of parameters available while activating only 3B routed parameters per token, keeping compute within a phone power budget while behaving like a much larger model.
More from Embodied
- Robotics startup Skild AI hits $100M ARR just 10 months after first deployment — deepakpathak · 2026-09-11
- McKinsey report calls humanoid robots a media distraction, sees just 2% of Physical AI market by 2045 — kscottz · 2026-09-11
- Simulated zebrafish and vision-equipped robot fish reveal how the body shapes brain circuits — DoctorJosh · 2026-09-11
- SpotHero + Tesla FSD gets you ~85% of a Waymo, says driver — cantrell · 2026-09-11
- Skild AI uses NVIDIA Physical AI to teach robots new tasks from a single video — nordicinst · 2026-09-11
- Play2Perfect (CoRL 2026): Play Pretraining Yields Precise Zero-Shot Robot Assembly — leto__jean · 2026-09-11