TokenSpeed Refactors KV Cache Architecture

AccBalanced · x · 2026-07-17

TokenSpeed detailed a scheduler refactor: moving away from a memory pool based on a Radix Tree—which became unsuitable as attention patterns grew complex—to a flat, block-based KV cache architecture.

The new implementation uses a single flat paged pool and provides heterogeneous views, making it easier to support different attention mechanisms. The post notes similar approaches in the community, such as vLLM's Jenga and LMDeploy's TurboMind.

This refactor roughly coincided with the release of TML's Inkling, allowing the new architecture to support Inkling from day one. The author also expressed commitment to continuing development with the open-source community.

Related event: Evolution and Refactoring of KV Cache Architecture in LLM Inference(2 posts)→

Original post →

More from Infra

Infra channel →