Running DeepSeek V4 Locally on Spark Hardware Hits ~95 tok/s

Rasmic · x · 2026-08-06

A developer shared a home lab setup test using Spark hardware. Data shows that running the DeepSeek V4 (0731) model requires just two Spark units, achieving an inference speed of approximately 95 tokens/sec.

The setup features a large context window and can even be driven by a single Spark device, enabling a fully local and private AI deployment of frontier-level models.

Original post →

More from Infra

Infra channel →