KV Cache Blending Boosts Prefill Speed 3x in Tests

maddie-lovelace · reddit · 2026-08-22

An engineer reports successful testing of KV cache blending for LLM inference. By splitting long prompts into chunks, generating caches in isolation, and concatenating them with overlap, the model maintains full needle-in-a-haystack retrieval and synthesis capabilities. The method boosted prefill speed by 3x on Ling3-tiny, hitting 1.3k tps at 256k tokens, matching Qwen 38-27B performance on a 5090 GPU.

Original post →

More from Infra

Infra channel →