v100-skinny: Open Source Kernels Push Qwen 27B Inference to 366 t/s on V100 GPUs
Simple_Library_2700 · reddit · 2026-08-12
Reddit user dnv2003 released v100-skinny, an open-source project featuring highly optimized inference kernels for Nvidia V100 GPUs (sm70 architecture). The kernels accelerate NVFP4 weight processing and offer nearly free speculative decoding.
In best-case multi-token prediction (MTP) extraction scenarios, the Qwen 27B model achieves an astonishing 366 t/s. For general code generation and MTP-friendly tasks, speeds hover around 200-240 t/s. The author notes that various caveats regarding these peak numbers are detailed in the GitHub repository.
More from Infra
- Musk: Starlink to carry >90% of internet traffic, nearly 11K satellites in orbit — DimaZeniuk · 2026-08-12
- Apple Silicon Virtualization Breakthrough: LLM Speeds Up 16x on macOS VMs — Scobleizer · 2026-08-12
- Xiaomi & Unitree Lock Global Shutter Sensor Capacity, Impacting US Robotics Scale-up — Scobleizer · 2026-08-12
- Bypassing Earthly Limits: Space Data Centers Emerge as AI Infrastructure Frontier — DavidLinthicum · 2026-08-12
- Muse Glimmer 30B Hits 3,323 tok/s on a Single NVIDIA GH200 — MaziyarPanahi · 2026-08-12
- NVIDIA Guide: Building a DGX Spark Dual-System Cluster via NVIDIA Sync — NVIDIA Developer · 2026-08-12