$2,800 rig of 8x Radeon Pro V620 (256GB VRAM) hits 3000+ t/s prefill via custom vLLM fork
_TheWolfOfWalmart_ · reddit · 2026-10-09
A Redditor built a $2,800 local inference rig with 8 used Radeon Pro V620 cards (RDNA2 cloud gaming GPUs, 32GB each, 256GB total VRAM):
- llama.cpp managed only 350-450 t/s prefill with poor concurrency, and vLLM didn't support the cards at all.
- He had Claude iteratively write and test custom RDNA2 kernels for a vLLM fork. Result: Qwen3.8-Flash-Next decodes at 60-100 t/s and prefills at 3000+ t/s—roughly 800% faster than before.
- Setup details: PP=4 (no tensor parallelism), routed experts W4A16 with BF16 elsewhere, MTP with 3-token drafting. Next targets: DeepSeek and GLM-5.3-Flash support.
More from coding & agent
- Procedural infinite city written 100% by Opus 5.5 runs in-browser via WebGPU — jason_mayes · 2026-10-09
- Voyager launches as an open harness for AI-driven video, graphics and games — jgooten · 2026-10-09
- TermGrade: 1k open-source executable terminal-agent RL environments with full training recipe — maximelabonne · 2026-10-09
- Step 5 Preview free in Cline for a week, beats Kimi K3 and GLM-5.3 on DeepSWE — StepFun_ai · 2026-10-09
- Open-source idea-to-launch prompt turns vague ideas into full go-to-market plans — Andrew0_0 · 2026-10-09
- Thrixel launches AI 3D creation engine that lets coding agents build editable, interactive 3D worlds — RanaHanocka · 2026-10-09