Local Dual-RTX Pro Inference Rig: 150 tok/s Decode, 10K tok/s Prefill, Full Build Notes

No_Run8812 · reddit · 2026-09-17

A Reddit user details building a local dual RTX Pro LLM workstation: 2.5 days of assembly including a PSU power fault and cooling rework, P2P + CUDA Graph + tensor parallelism setup, GPUs capped at 500W. It hits 150 tok/s decode and 10K tok/s prefill, running Qwen 3.8 flash next 8-bit and DeepSeek V4 flash 0731. They find the M3 Ultra 512GB too slow for inference by comparison, prefer DeepSeek's 1M context with room for 4 concurrent requests, and have replaced subscriptions with openclaw + opencode + Open WebUI + Tailscale.

Original post →

More from coding & agent

coding & agent channel →