Software magic: 4x V100s match RTX 5090 running Qwen 3.8 NVFP4
Simple_Library_2700 · reddit · 2026-08-19
A developer achieved performance parity between four Tesla V100s (2017) and an RTX 5090 when running Qwen 3.8 NVFP4 by writing a custom QPN kernel. Although V100s lack native FP4/FP8 support, the kernel translates fragments on the fly to FP16 for Volta's Tensor Cores. In single-request decoding, the V100 setup reached 219.1 tok/s, slightly edging out the 5090's 214.7 tok/s by verifying more tokens per round (5.89 vs 4.27).
More from Infra
- DSCO Router Launches Unified Gateway for Multi-Model Routing with BYOK Support — arthurcolle · 2026-08-24
- Open Source RobotSoul: Persistent Identity for Agents After Context Resets — robauto-dot-ai · 2026-08-24
- Offloading MoE models to RAM causes slow prefill speeds — former_farmer · 2026-08-24
- Etched Raises $1B Led by Jane Street to Validate Architecture-Agnostic AI Chips — TheTuringPost · 2026-08-24
- ConvRot Quant joins llama-cpp: Q6 accuracy nears Q8 quality — giveen · 2026-08-24
- LifeOS: A Local, Voice-Driven Personal Organizer — Extension-Bid-639 · 2026-08-24