Dev Runs Uncensored DeepSeek Locally Using Dual Sparks at 40 tok/s
PMinervini · x · 2026-08-09
Developer Teknium shared their local deployment experience: by connecting two Spark units with a single cable, they ran an 'abliterated' (uncensored) DeepSeek v4 flash 0731 model without dspark, achieving completely private inference at around 40 tokens/s.
Other developers subsequently inquired about the exact method used for abliteration, speculating whether it utilized ds4's direction steering feature.
Related event: Dual Spark Runs Abliterated DeepSeek Locally at 40 tok/s(3 posts)→
More from Infra
- Cursor's 'Git at Any Scale' Praised: Stateless Storage to Rewrite Web Infrastructure — thesephist · 2026-08-24
- AI Performance Engineering resource list v2 covers everything from CUDA to MoE serving — AccBalanced · 2026-08-24
- Semiconductor engineers now more prestigious than doctors in South Korea — SuB8u · 2026-08-24
- The VendorOps dilemma in modern-day programming — vboykis · 2026-08-24
- Local development is fast and controllable, why rely solely on the cloud? — vboykis · 2026-08-24
- Trained on PrimeIntellect, rollouts rendered on Modal — willcb · 2026-08-24