A vendor-agnostic Vulkan backend cuts edge inference latency from 30 ms to 3 ms
ppchaos · reddit · 2026-07-29
A production edge setup for a video editing tool shows how to run ML inference across heterogeneous hardware without relying on CUDA.
The team uses ncnn’s Vulkan backend so the same stack works on NVIDIA, AMD, Intel integrated graphics, and Apple Silicon. On an RTX 4070 in fp16, ArcFace R50 drops from 30 ms on ONNX CPU to 3 ms on ncnn Vulkan, while SCRFD falls from 25 ms to 2.5 ms. Model storage also shrinks, with ArcFace going from 174 MB ONNX fp32 to 87 MB in ncnn fp16 weight storage. The main argument is not just speed: Vulkan exists on the machines they ship to, so users do not need vendor-specific runtimes or installs.
More from Infra
- RTX 5090 cards are selling for 40,000 SEK plus VAT in another GPU price spike — wandedob · 2026-07-29
- Helion lands in Hugging Face Kernels and beats PyTorch SDPA on H100 by 1.17× — RisingSayak · 2026-07-29
- Extropic signs a $75 million Commerce Department LOI to scale thermodynamic AI chips — beffjezos · 2026-07-29
- On two L40S GPUs, vLLM tensor parallelism beat pipeline parallelism in a Qwen3.6-35B FP8 benchmark — amiitk · 2026-07-29
- SK hynix reportedly signs long-term contracts with about 10 major customers — dejavucoder · 2026-07-29
- Shanghai Aishengna is said to be manufacturing DUV lithography tools — zephyr_z9 · 2026-07-29