DeepSeek 4 Flash runs all day on Spark: zero crashes, 98% cache hit, 40 tok/s
jasonkneen · x · 2026-09-09
Developer Jason Kneen ran DeepSeek 4 Flash on a Spark all day for tool unification and doc generation: zero crashes, 98% cache hit rate and 40 tok/s with RAM nearly maxed. He's setting up a second Spark, showing consumer devices can handle long-running local agent workloads.
More from Infra
- vLLM's Hybrid HiSparse keeps decoding past HBM limits: 19-25 vs 5-6 concurrent 1M-context requests — vllm_project · 2026-09-09
- $22.5M of Compute in One Week: The New Price Tag for Math Breakthroughs — vykthur · 2026-09-09
- $1M sounds like a lot — it's just 10.5 minutes of OpenAI's compute spend — wordgrammer · 2026-09-09
- SageMaker Feature Store adds UpdateRecord for atomic feature-level writes, no more read-modify-write — AWS ML Blog · 2026-09-09
- Modular shows unified compute layer: TPU v6e and d-Matrix Corsair brought up in days — carrycooldude · 2026-09-09
- LM Studio ships Bionic 1.1.2 with faster sessions, better screen reader support, Linux builds — mattturck · 2026-09-09