Can Dual Xeon Run Large Models?
dbinnunE3 · reddit · 2026-07-18
The author currently owns a Nimo Strix Halo machine for smaller local models and internal scripts, and is evaluating whether a Dell T7920 dual Xeon server can serve as an LLM inference rig.
The core question: with 1.5TB DDR4, dual CPUs, and two PCIe slots for GPUs, can it run larger open-source models via layer splitting/weight loading? Furthermore, would adding one or two GPUs create a "good enough but interesting" LLM workstation? The author already uses tools like llama-swap, llama.cpp, OpenWebUI, MCP/skills, but is unfamiliar with the PP/TG performance and deployment for dual-socket servers and is seeking advice and guidelines.
More from Infra
- China’s AI arms race is increasingly defined by chips, data centers, and open models — BenBajarin · 2026-07-22
- Agent search bottlenecks are now about variance, not raw latency — rohanpaul_ai · 2026-07-22
- Gavin Baker argues Nvidia may be one of open source AI’s biggest supporters — GavinSBaker · 2026-07-22
- AI Power Demand Exposes US Energy Gap, Urging Shift from Scarcity to Abundance — bradneuberg · 2026-07-22
- Gavin Baker says Nvidia’s $630B figure would be system revenue, not all Nvidia’s — GavinSBaker · 2026-07-22
- A Firecracker-based platform says it can host 6,000 AI agents on one 256 GB server — maritime_sh · 2026-07-22