1Cat-vLLM fork revives decade-old V100s: Qwen3.6-35B benchmark numbers shared
Miserable-Dare5090 · reddit · 2026-09-26
For users still holding V100 cards, the author points to 1Cat-vLLM, a vLLM fork enabling optimized serving for Volta cards. They share raw llama-benchy numbers running Qwen3.6-35B, comparing a Strix Halo box with a heavily optimized llama.cpp fork (pwilkin) against the V100 + 1Cat setup—not apples to apples, but impressive for 10-year-old GPUs—and invite further optimization suggestions.
More from Infra
- Running local AI on a MacBook Pro: fan noise fixed by switching to automatic power mode — walkingriver · 2026-09-26
- Benchmark maker says no Ascend version — models would hill-climb it; TPU/Trainium/AMD better — xeophon · 2026-09-26
- FT: Oracle Must Pay Data Center Investors Even If Sites Never Get Power — SumitGup · 2026-09-26
- China Telecom's Xing4.0 trends on HF: 29B MoE trained entirely on Ascend 910C — AdinaYakup · 2026-09-26
- Data Center Resistance Is Growing in Africa: Scholar Digs Into the Backlash — ChinasaTOkolo · 2026-09-26
- Citi: monthly AI token usage has grown 31% MoM on average since early 2025 — Beth_Kindig · 2026-09-26