Looking for a cheap DGX Spark alternative to serve 15-20GB of embedding models locally
fuse1921 · reddit · 2026-09-02
A user with a 4x3090 local cluster is asking for hardware advice: with PCIe lanes exhausted on the main box, they don't want to build a whole new machine just to serve memory-system LLMs (embeddings, rerankers) needing 15-20GB of VRAM, currently running them on a gaming PC's 4090.
They want a turnkey, low-power, local (no VPS) inference box like NVIDIA DGX Spark but cheaper ($1,000), and are worried pure CPU would be too slow. The thread solicits community recommendations.
More from Infra
- Tencent open-sources CubeSandbox v0.7.0, keeping thousands of agents alive through node failures — SucceededMind · 2026-09-03
- Magnitude open-source inference server swaps coding agents to free local models automatically — nickbaumann_ · 2026-09-03
- MiniMax H3 video generation runs fully local on a single RTX 5060 Ti 16GB — apoke890 · 2026-09-03
- With 98% Cache Hits in Coding, 400-600 t/s Prefill Already Hits Diminishing Returns — nomorebuttsplz · 2026-09-03
- Dell COO: inference tokens to grow 87x to 3,600 quadrillion by 2030 — Beth_Kindig · 2026-09-03
- Beating hipBLASLt on a gaming GPU: a GEMM optimization deep dive — Moist_Weird_42067 · 2026-09-03