Reka EdgeQ VLM Runs Natively on Snapdragon 8 Elite NPU With 0.73s First Token
RekaAILabs · x · 2026-09-23
Reka launched Reka EdgeQ, an on-device VLM running natively on Qualcomm Snapdragon 8 Elite's Hexagon NPU for wearables and phones: 0.73s time-to-first-token on images, +34 pts over Gemma 4 E4B on MLVU video benchmarks, and 6.9 mWh per inference (3x lower heat overhead). They custom-tuned a ConvNeXt V2 vision encoder/decoder for the NPU, keeping the GPU idle to avoid thermal throttling.
More from Infra
- Presenting Undersea Internet Cables at the United Nations — EricTopol · 2026-09-23
- NVIDIA's AI Factory Insider Ep 5: How CUDA Powers Specialized AI Across Industries — NVIDIA Developer · 2026-09-23
- vLLM adds hardware-agnostic layers: within 3.4% of native throughput on H100 while keeping portability — PyTorch · 2026-09-23
- 421M-Parameter Laya Model Plays Flappy Bird on CPU via OpenVINO INT8 — simpleuserhere · 2026-09-23
- How to run Qwen3.8-27B with 160k context on a 16GB AMD card: full config — According_Study_162 · 2026-09-23
- MLX-Serve v26.9.5 lands with Qwen-Image 2.1 and 4-way MTP streams at up to 122 tok/s on M4 Max — TheMoonMidas · 2026-09-23