512GB DDR4 + 2x RTX 3090: What Local Models Should You Run on This Setup?
ArtifartX · reddit · 2026-09-12
A Reddit user is building a server dedicated to local LLM inference including agentic coding: 512GB DDR4-2666 with 2x RTX 3090 (48GB VRAM), expandable to 7 GPUs on PCIe 4.0 x16. They ask whether to run several smaller models or one large one, citing a setup running Qwen 3.8-Flash-Next at 9.12 tok/s decode on dual 3090s. The thread covers RAM/VRAM tradeoffs and quantization practicality.
More from Infra
- USD.AI lends against tokenized GPUs; largest loan grew from $620K to $98.1M in a year — antavedissian · 2026-09-12
- USD.AI: deposits fled fast but loan book kept growing; $557M TVL, $322M pipeline — antavedissian · 2026-09-12
- Nari Labs launches 40ms speech-to-text at $0.06/hr, claims 9x cheaper than Gemini Transcribe — alexcovo_eth · 2026-09-12
- NVIDIA announces TensorRT Model Connect for faster video-to-voice builds — NVIDIAAI · 2026-09-12
- Hugging Face Kernels adds Helion support for shipping autotuned ML kernels — PyTorch · 2026-09-12
- Sleeping models in Rust: reclaiming 79GB of idle VRAM with sub-200ms wake — ahstanin · 2026-09-12