Can 2x RTX 3090 Plus 512GB DDR5 Reach 15 t/s on Large Local LLMs? Redditor Asks
levoniust · reddit · 2026-09-11
A Redditor seeks upgrade advice for local LLM inference: with 512GB DDR5-5600 UDIMM, two Core Ultra 7 265K systems, and 2x RTX 3090 (48GB VRAM), running DeepSeek-R1-0528 Q3 yields only 2 tokens/s. They ask whether Threadripper/EPYC/Xeon platforms could hit 15+ t/s while reusing the GPUs, requesting specific CPU/motherboard picks and firsthand benchmarks.
More from Infra
- Bartowski Details New Per-Tensor Layout Maps for GGUF Quantization in llama.cpp — bartowski1182 · 2026-09-11
- antirez runs DeepSeek v4.1 Flash locally on a 128GB M5 Max, SSD streaming surprisingly fast — antirez · 2026-09-11
- OpenAI could 7x its training compute tomorrow: why open-source models still trail by one generation — soumitrashukla9 · 2026-09-11
- DOJ scrutinizes Nvidia's ~$20B Groq licensing deal over merger-review evasion — eyishazyer · 2026-09-11
- Persimmon Built on NVIDIA's 550B Nemotron 3 Ultra with Thousands of Blackwell GPUs — niloofar_mire · 2026-09-11
- NVIDIA details EPD disaggregation: up to 5x faster TTFT and 7x faster responses for multimodal serving — NVIDIAAI · 2026-09-11