Running DeepSeek Locally on Dual Sparks: 40 tok/s Uncensored Inference

Teknium · x · 2026-08-09

A developer shared a hands-on experience running the uncensored (abliterated) DeepSeek V4 Flash 0731 locally using two Spark devices connected with a single cable.

This setup achieves an inference speed of around 40 tokens/second without using dSpark, enabling completely private and uncensored local inference. Teknium confirmed the setup works great in Hermes Agent.

Related event: Dual Spark Runs Abliterated DeepSeek Locally at 40 tok/s(3 posts)→

Original post →

More from Infra

Infra channel →