Qualcomm says NPU latency beats GPU and CPU for on-device AI agents
qdrant_engine · x · 2026-07-28
Qualcomm’s Alan Zhu argued that the most underestimated benefit of on-device AI is latency, not just cost or privacy.
At Vector Space Day SF, he ran the same agentic search workload on three compute paths:
- NPU: responded instantly
- GPU: took about 12 seconds
- CPU: took about 30 seconds, then slowed further as the device heated up
He said the NPU sustained 90 tokens/sec for the full session at 37°C. The point: fast NPU prefill makes it practical to search thousands of local files and return a cited answer in seconds, fully on-device and without network dependence.
More from coding & agent
- x402 Builder Codes add revshare for apps and agents routing demand — kleffew94 · 2026-07-28
- Claude Code creator says startups should chase the model capabilities nobody has productized yet — ycombinator · 2026-07-28
- Different coding agents can produce very different estimates from the same model spec — JessicaHullman · 2026-07-28
- Progressive context extension may suit linear-attention hybrids better than short-SWA models — stochasticchasm · 2026-07-28
- Codex automation checked SF apartment listings hourly and won the lease first — nickbaumann_ · 2026-07-28
- Renting GPUs and open-weight models cut one AI bill from $1.2M to $100K — kimmonismus · 2026-07-28