10-Person Startup Builds 4x AMD R9700 Local Inference Box, Posts Real Numbers
sayamss · reddit · 2026-10-11
A developer shared a local inference build for a 10-person startup: Threadripper 9970x, 128GB DDR5 ECC, and 4x AMD Radeon AI Pro R9700 (128GB total VRAM), with GPUs undervolted under 210W each on a 1600W PSU. Serving Qwen 3.8 Next flash/27b and DSV4 Flash via a Radiance fork.
- 16 concurrent sessions: Qwen 3.8 27b MXFP4 hits 6.3-6.8k tok/s aggregate prefill, 80-900 t/s aggregate decode depending on context
- Qwen 3.8 Next flash: 365 tok/s aggregate at 128k context/user, 538.5 tok/s at 48k
- BF16 KV cache throughout
A useful reference config and benchmark for teams considering self-hosted inference.
More from Infra
- Zeeg: persistent VMs for agents are wrong, ephemeral sandboxes are the present — zeeg · 2026-10-11
- OpenRouter Processes ~$1.5B Annualized Inference, Anthropic Takes 36% of Spend — deedydas · 2026-10-11
- AMD Reportedly Raises GDDR6 Prices for Board Partners Amid GPU Price Pressure — chemist_slime · 2026-10-11
- M5 Ultra Mac Studio: is 80-core worth +$1200 and 10 extra weeks over 64-core? — Porespellar · 2026-10-11
- Virginia animal shelter says nearby data center noise is terrifying its rescue dogs — Polymarket · 2026-10-11
- Running Qwen3.8-27B on an RTX 5090 with vLLM: NVFP4 works, FP8 doesn't — BitGreen1270 · 2026-10-11