1Cat-vLLM fork revives decade-old V100s: Qwen3.6-35B benchmark numbers shared

Miserable-Dare5090 · reddit · 2026-09-26

For users still holding V100 cards, the author points to 1Cat-vLLM, a vLLM fork enabling optimized serving for Volta cards. They share raw llama-benchy numbers running Qwen3.6-35B, comparing a Strix Halo box with a heavily optimized llama.cpp fork (pwilkin) against the V100 + 1Cat setup—not apples to apples, but impressive for 10-year-old GPUs—and invite further optimization suggestions.

Original post →

More from Infra

Infra channel →