Running DeepSeek Locally on Dual Sparks: 40 tok/s Uncensored Inference
Teknium · x · 2026-08-09
A developer shared a hands-on experience running the uncensored (abliterated) DeepSeek V4 Flash 0731 locally using two Spark devices connected with a single cable.
This setup achieves an inference speed of around 40 tokens/second without using dSpark, enabling completely private and uncensored local inference. Teknium confirmed the setup works great in Hermes Agent.
Related event: Dual Spark Runs Abliterated DeepSeek Locally at 40 tok/s(3 posts)→
More from Infra
- Cursor's 'Git at Any Scale' Praised: Stateless Storage to Rewrite Web Infrastructure — thesephist · 2026-08-24
- AI Performance Engineering resource list v2 covers everything from CUDA to MoE serving — AccBalanced · 2026-08-24
- Semiconductor engineers now more prestigious than doctors in South Korea — SuB8u · 2026-08-24
- The VendorOps dilemma in modern-day programming — vboykis · 2026-08-24
- Local development is fast and controllable, why rely solely on the cloud? — vboykis · 2026-08-24
- Trained on PrimeIntellect, rollouts rendered on Modal — willcb · 2026-08-24