Looking for a cheap DGX Spark alternative to serve 15-20GB of embedding models locally

fuse1921 · reddit · 2026-09-02

A user with a 4x3090 local cluster is asking for hardware advice: with PCIe lanes exhausted on the main box, they don't want to build a whole new machine just to serve memory-system LLMs (embeddings, rerankers) needing 15-20GB of VRAM, currently running them on a gaming PC's 4090.

They want a turnkey, low-power, local (no VPS) inference box like NVIDIA DGX Spark but cheaper ($1,000), and are worried pure CPU would be too slow. The thread solicits community recommendations.

Original post →

More from Infra

Infra channel →