MacBook + DGX Spark: Testing Heterogeneous Inference and KV Cache Shipping
HankYeomans · x · 2026-08-11
A developer conducted a heterogeneous inference experiment: using DGX Spark for prefill and an M4 Max MacBook Pro for decode, leveraging each hardware's strengths.
- Core Concept: Utilizing Spark's fast 1.5-2.3k tok/s prefill speed to overcome Mac's long context bottleneck, then shipping the generated KV cache file to the Mac for decoding.
- Transfer Projections: DeepSeek's cache is 86 KB/token. Over a 10GbE wired connection, shipping a 128k context (5.8 GB) takes 5s; 500k takes 20s; 1M takes 42s. Wireless transfer times are significantly longer.
- Status: WiFi testing is underway; wired metrics are theoretical projections.
More from Infra
- ComfyUI Plugin Optimizes LoRA Loading, Slashing MiniMax-H3 VRAM by 38GB — marres · 2026-08-11
- Can a Single NVIDIA DGX Replace All Your AI Subscriptions? — jackedAJ · 2026-08-11
- Rumored 50-Series Super Bumps VRAM: 5070 Ti to 24GB — PROfil_Official · 2026-08-11
- RTX 3060 Tests: How GPU Memory Allocation Shifts LLM Execution Strategy — Abhishekcur · 2026-08-11
- CXMT 17nm DDR5 Yield Exceeds 90%, But US PC Makers Restrict Procurement — teortaxesTex · 2026-08-11
- Open-Source Models Cut Inference Costs 8x, Compute Becomes New Bottleneck — latticecut · 2026-08-11