Running DeepSeek V4 Locally on Spark Hardware Hits ~95 tok/s
Rasmic · x · 2026-08-06
A developer shared a home lab setup test using Spark hardware. Data shows that running the DeepSeek V4 (0731) model requires just two Spark units, achieving an inference speed of approximately 95 tokens/sec.
The setup features a large context window and can even be driven by a single Spark device, enabling a fully local and private AI deployment of frontier-level models.
More from Infra
- Google Cloud's Filestore Migrates to Colossus, Decoupling Capacity from IOPS — rseroter · 2026-08-06
- Testing 8x DGX Spark Nodes in Open World Multi-Agent Setup — NVIDIAAI · 2026-08-06
- Open-Source Benchmarks: RTX 5090 LLM Quants and 8GB VRAM Agentic Scores — max_paperclips · 2026-08-06
- NVIDIA Discusses Building Secure Enterprise AI with Proprietary Data — nvidia · 2026-08-06
- Chorus: Open-Source Pre-trained Model Library Enables Fast CPU Inference Without GPUs — jmschreiber91 · 2026-08-06
- Local Deployment: Running an NVIDIA and AMD GPU Together for Different Models — Curious-Pen5547 · 2026-08-06