Can 2x RTX 3090 Plus 512GB DDR5 Reach 15 t/s on Large Local LLMs? Redditor Asks

levoniust · reddit · 2026-09-11

A Redditor seeks upgrade advice for local LLM inference: with 512GB DDR5-5600 UDIMM, two Core Ultra 7 265K systems, and 2x RTX 3090 (48GB VRAM), running DeepSeek-R1-0528 Q3 yields only 2 tokens/s. They ask whether Threadripper/EPYC/Xeon platforms could hit 15+ t/s while reusing the GPUs, requesting specific CPU/motherboard picks and firsthand benchmarks.

Original post →

More from Infra

Infra channel →