512GB DDR4 + 2x RTX 3090: What Local Models Should You Run on This Setup?

ArtifartX · reddit · 2026-09-12

A Reddit user is building a server dedicated to local LLM inference including agentic coding: 512GB DDR4-2666 with 2x RTX 3090 (48GB VRAM), expandable to 7 GPUs on PCIe 4.0 x16. They ask whether to run several smaller models or one large one, citing a setup running Qwen 3.8-Flash-Next at 9.12 tok/s decode on dual 3090s. The thread covers RAM/VRAM tradeoffs and quantization practicality.

Original post →

More from Infra

Infra channel →