KV Cache Blending Boosts Prefill Speed 3x in Tests
maddie-lovelace · reddit · 2026-08-22
An engineer reports successful testing of KV cache blending for LLM inference. By splitting long prompts into chunks, generating caches in isolation, and concatenating them with overlap, the model maintains full needle-in-a-haystack retrieval and synthesis capabilities. The method boosted prefill speed by 3x on Ling3-tiny, hitting 1.3k tps at 256k tokens, matching Qwen 38-27B performance on a 5090 GPU.
More from Infra
- H100 Shortage Driven by Power, Cooling, and Talent, Not Chip Fabrication — ingliguori · 2026-08-22
- Parsewave's work suggests a shift towards high-quality synthetic data in AI training — trashnash007 · 2026-08-22
- GLM 5.2 runs at 35t/s via DwarfStar mixed RAM/VRAM inference — antirez · 2026-08-22
- Data center pause may push AI jobs overseas amidst power bottleneck fears — Dan_Jeffries1 · 2026-08-22
- GLM-5.3 hits 21.4x speedup on RTX PRO 6000 via Kimi-Linear Decode kernel — teortaxesTex · 2026-08-22
- Why I'm traveling to India to build a dual-RTX 3090 rig for local LLMs — Ubunta · 2026-08-22