llama.cpp NVFP4 Dot Product Lookup Optimization Boosts ARM Speed by ~5x

pmttyji · reddit · 2026-07-06

A llama.cpp community PR extends UE4M3 lookup table optimization to NVFP4 dot product operations on ARM, reusing existing lookup infrastructure to align ARM implementation with x86. Benchmarks show that the Qwen3.5-4B-NVFP4 model on 4 threads saw pp512 performance jump from 1.89 t/s to 9.97 t/s, achieving a roughly 5x speedup.

Original post →

More from Infra

Infra channel →