Dev Runs Uncensored DeepSeek Locally Using Dual Sparks at 40 tok/s

PMinervini · x · 2026-08-09

Developer Teknium shared their local deployment experience: by connecting two Spark units with a single cable, they ran an 'abliterated' (uncensored) DeepSeek v4 flash 0731 model without dspark, achieving completely private inference at around 40 tokens/s.

Other developers subsequently inquired about the exact method used for abliteration, speculating whether it utilized ds4's direction steering feature.

Related event: Dual Spark Runs Abliterated DeepSeek Locally at 40 tok/s(3 posts)→

Original post →

More from Infra

Infra channel →