$3000 home server with 128GB VRAM runs Qwen3.8-next at 1.3k tps prefill, 70 tps code
Thin_Pollution8843 · reddit · 2026-09-14
- A Redditor finished a home inference server costing $3000 (SSD excluded): 128GB VRAM + 256GB DDR4 RAM.
- Build: 4x RTX V620 ($1400), 256GB DDR4 RDIMM 2666 ($610), Huanandzhi D12D board ($410), EPYC 7452 ($170), ASRock 1600W PSU ($220), case/fans/misc $200. An earlier Lenovo P620 attempt was returned over proprietary parts.
- Power draw is real: 700–900W prefill, 500–600W decode running Qwen3.8-next-flash Autoround W4A16.
- Results: with a vllm fork + MTP-2 at 128k+ context, 1.3k tps prefill and 70 tps code / 60 tps prose generation. The author was initially disappointed with Qwen3.8-27b speeds but is satisfied with Qwen3.8-next, hoping for more optimizations.
More from Infra
- Used RTX 5090 listed at £3,900 (~$5,200), more than double its MSRP — julianharris · 2026-09-14
- Why Amazon and Microsoft Are Taking Communities' Side Against Utilities — pstAsiatech · 2026-09-14
- Redditor builds dual AI workstation with 2x RTX 3090 and 4x Tesla P100, asks how to optimize — FearFactory2904 · 2026-09-14
- agi-memory: SQLite-only persistent memory MCP server for coding assistants, 32MB RAM — Rude_Gate7599 · 2026-09-14
- Top 10 foundry revenue hits $53.5B in Q2, TSMC holds 72.5% share — Beth_Kindig · 2026-09-14
- Qwen3.8 27B INT4 With 144K Context Runs on a Single RTX 3090 via vLLM — Altruistic_Heat_9531 · 2026-09-14