v100-skinny: Open Source Kernels Push Qwen 27B Inference to 366 t/s on V100 GPUs

Simple_Library_2700 · reddit · 2026-08-12

Reddit user dnv2003 released v100-skinny, an open-source project featuring highly optimized inference kernels for Nvidia V100 GPUs (sm70 architecture). The kernels accelerate NVFP4 weight processing and offer nearly free speculative decoding.

In best-case multi-token prediction (MTP) extraction scenarios, the Qwen 27B model achieves an astonishing 366 t/s. For general code generation and MTP-friendly tasks, speeds hover around 200-240 t/s. The author notes that various caveats regarding these peak numbers are detailed in the GitHub repository.

Original post →

More from Infra

Infra channel →