Can Dual Xeon Run Large Models?

dbinnunE3 · reddit · 2026-07-18

The author currently owns a Nimo Strix Halo machine for smaller local models and internal scripts, and is evaluating whether a Dell T7920 dual Xeon server can serve as an LLM inference rig.

The core question: with 1.5TB DDR4, dual CPUs, and two PCIe slots for GPUs, can it run larger open-source models via layer splitting/weight loading? Furthermore, would adding one or two GPUs create a "good enough but interesting" LLM workstation? The author already uses tools like llama-swap, llama.cpp, OpenWebUI, MCP/skills, but is unfamiliar with the PP/TG performance and deployment for dual-socket servers and is seeking advice and guidelines.

Original post →

More from Infra

Infra channel →