"Sharding Is the Transfer": MindLab Breaks 2M-Context VRAM Wall with 2D KV Resharding

青稞AI · wechat · 2026-08-30

After launching Macaron-V1, MindLab found their 2M-context service hitting VRAM saturation at scale: each rank had only 1.17M tokens of KV capacity left, and just 20–40 concurrent requests triggered queuing, with 70% of request time spent waiting. Rather than truncating or summarizing history (they refuse to compromise agent memory quality), they optimized along three axes:

The core innovation is Page-Level 2D Resharding: prefill shards by layer, decode by page — orthogonal dimensions. The naive approach (transfer whole KV, then reshard) wastes 4x bandwidth. Since RDMA (Mooncake) and DCP both operate on pages, they fused sharding into the transfer itself: each rank pre-filters by page number and data lands in its final slot, cutting traffic to 1/4. Consensus protocols (size-broadcast, prefix-length agreement) prevent deadlocks.

Production results: same hardware went from 20–40 concurrent queuing to 100+ stable, peak decode throughput from 800 to 1800 tokens/s.

Original post →

More from Infra

Infra channel →