Software magic: 4x V100s match RTX 5090 running Qwen 3.8 NVFP4

Simple_Library_2700 · reddit · 2026-08-19

A developer achieved performance parity between four Tesla V100s (2017) and an RTX 5090 when running Qwen 3.8 NVFP4 by writing a custom QPN kernel. Although V100s lack native FP4/FP8 support, the kernel translates fragments on the fly to FP16 for Volta's Tensor Cores. In single-request decoding, the V100 setup reached 219.1 tok/s, slightly edging out the 5090's 214.7 tok/s by verifying more tokens per round (5.89 vs 4.27).

Original post →

More from Infra

Infra channel →