"Sharding Is the Transfer": MindLab Breaks 2M-Context VRAM Wall with 2D KV Resharding
青稞AI · wechat · 2026-08-30
After launching Macaron-V1, MindLab found their 2M-context service hitting VRAM saturation at scale: each rank had only 1.17M tokens of KV capacity left, and just 20–40 concurrent requests triggered queuing, with 70% of request time spent waiting. Rather than truncating or summarizing history (they refuse to compromise agent memory quality), they optimized along three axes:
- HiCache tiered caching: KV cache across GPU/CPU/storage tiers; repeated history prefixes skip re-prefill, keeping hot-request TTFT under 5s.
- EAGLE speculative decoding: reuses GLM-5.2's built-in MTP layer as draft head — TPOT drops from 30.8ms to 8.6ms (3.6x).
- PD disaggregation + CP + DCP: prefill and decode deployed separately, split by layers (78 layers / 8 ranks) and by pages (64 tokens / page, 4 ranks) respectively, scaling decode-side KV capacity linearly with TPOT overhead of only 5–15%.
The core innovation is Page-Level 2D Resharding: prefill shards by layer, decode by page — orthogonal dimensions. The naive approach (transfer whole KV, then reshard) wastes 4x bandwidth. Since RDMA (Mooncake) and DCP both operate on pages, they fused sharding into the transfer itself: each rank pre-filters by page number and data lands in its final slot, cutting traffic to 1/4. Consensus protocols (size-broadcast, prefix-length agreement) prevent deadlocks.
Production results: same hardware went from 20–40 concurrent queuing to 100+ stable, peak decode throughput from 800 to 1800 tokens/s.
More from Infra
- OpenAI's Jalapeño Chip Beats Nvidia GB200 in Efficiency — Beth_Kindig · 2026-08-30
- Developer Dumps Qualcomm for Rockchip to Ship Edge AI Devices Faster — kscottz · 2026-08-30
- Empire of AI's datacenter water figure is off by ~4,500x: liters vs cubic meters — altryne · 2026-08-30
- Community fork enables MiniMax H3 on dual GPUs with live preview support — karma3u · 2026-08-30
- Own your harness, and if possible, own the model layer too — omarsar0 · 2026-08-30
- GLM 5.3 Flash Inference Extremely Slow on Apple Silicon — CentrifugalMalaise · 2026-08-30