Optimizing DeepSeek V4 on Mac Studio: 12x Speedup & KV Cache Tricks
Adrian_Galilea · reddit · 2026-08-19
Author optimized DeepSeek V4 Flash on a Mac Studio M3 Ultra (512GB), reducing response latency from 6-20s to 1.6s.
Kernel Optimizations (+21% prefill speed)
Targeted the "lightning indexer" bottleneck in sparse attention with three PRs:
- Threadgroup-tiled scorer
- Register-blocked K-resident scorer
- Streaming top-512 replacing bitonic sort
KV Cache Prewarming (Universal 10x gain)
- State Synchronization: Clients must replay the exact model output verbatim. Mismatched prefixes cause cache misses and full re-prefill per turn.
- Prewarming: Send requests with maxtokens: 0. This forces the engine to prefill and stop exactly at the prompt, allowing subsequent real requests to extend the cache with zero latency.
More from Infra
- Ling-3.0-tiny Runs 128K Context on $249 8GB Orin Nano — Puzzleheaded_Base302 · 2026-08-19
- Apple's Foundation Model Framework: Hybrid AI Routing with Dynamic Profiles — Scobleizer · 2026-08-19
- Docling Graph turns documents into queryable knowledge graphs using Pydantic — techNmak · 2026-08-19
- CoreWeave hits $2.6B quarterly revenue in just 25 quarters, a milestone AWS took 40 to reach — FinanceYF5 · 2026-08-19
- Accelerating Feature Engineering with GPU: A Practical Guide to Target Encoding — pandeyparul · 2026-08-19
- Train AI Models Locally via Desktop App, Connect Claude Code with One Command — Saboo_Shubham_ · 2026-08-19