Local Dual-RTX Pro Inference Rig: 150 tok/s Decode, 10K tok/s Prefill, Full Build Notes
No_Run8812 · reddit · 2026-09-17
A Reddit user details building a local dual RTX Pro LLM workstation: 2.5 days of assembly including a PSU power fault and cooling rework, P2P + CUDA Graph + tensor parallelism setup, GPUs capped at 500W. It hits 150 tok/s decode and 10K tok/s prefill, running Qwen 3.8 flash next 8-bit and DeepSeek V4 flash 0731. They find the M3 Ultra 512GB too slow for inference by comparison, prefer DeepSeek's 1M context with room for 4 concurrent requests, and have replaced subscriptions with openclaw + opencode + Open WebUI + Tailscale.
More from coding & agent
- Routing simple requests to a cheap model made our total LLM bill worse — Massive_Tell_4276 · 2026-09-17
- Dev loves Codex app's /side chat so much he bound it to ⌘S — alex_frantic · 2026-09-17
- From personal second brain to team brain: one table, labels, and MCP access — Cole Medin · 2026-09-17
- AI SDK Adds evaluate Method, Welcomes Typesafe AI's New Jev Model — cramforce · 2026-09-17
- BotBell MCP lets AI assistants push notifications to iPhone and Mac — modelcontextprotocol · 2026-09-17
- Hüttentouren MCP adds live hut-to-hut hiking availability in the Alps — modelcontextprotocol · 2026-09-17