Dual DGX Spark Setup Hits 40 tok/s on DeepSeek Locally

Teknium · x · 2026-08-09

Developer Teknium shared real-world benchmarks for local AI deployment using NVIDIA DGX Spark. By connecting two Spark units with a single cable, he achieved around 40 tok/s running an uncensored DeepSeek model without extra acceleration frameworks. This enables completely private, local inference. He also expressed hope that the upcoming Spark 2 will feature 512GB of memory.

Related event: Dual Spark Runs Abliterated DeepSeek Locally at 40 tok/s(3 posts)→

Original post →

More from Infra

Infra channel →