SparseEngine: sparse-first inference engine delivers 10x throughput with KV eviction

Jitai Hao · hf · 2026-10-09

SparseEngine is a ground-up sparse-first inference engine supporting 15 sparse attention methods via a shared lifecycle contract, with Chain Cache for cross-request KV-eviction state resumption and controllable prefix-cache pruning. It delivers over 10x throughput with KV eviction, 2.5x faster decoding than vLLM at matched concurrency, and 2x end-to-end speedup on agent benchmarks. Code is open source.

Original post →

More from Infra

Infra channel →