SparseEngine: sparse-first inference engine delivers 10x throughput with KV eviction
Jitai Hao · hf · 2026-10-09
SparseEngine is a ground-up sparse-first inference engine supporting 15 sparse attention methods via a shared lifecycle contract, with Chain Cache for cross-request KV-eviction state resumption and controllable prefix-cache pruning. It delivers over 10x throughput with KV eviction, 2.5x faster decoding than vLLM at matched concurrency, and 2x end-to-end speedup on agent benchmarks. Code is open source.
More from Infra
- 27B model on a single RTX 4090: 262K context at ~130 tok/s with NInfer — Distinct-Pie2389 · 2026-10-09
- boat spins up 250 agent sandbox VMs for 10 cents: full Ubuntu boxes at $20/mo — RexDouglass · 2026-10-09
- A cache hit is not free: inside Triton's compilation cache and hidden costs — Mahmoud_Zalt · 2026-10-09
- AI agents could spawn history's largest bureaucracy, where machines create work for machines — brucemacv · 2026-10-09
- Bittensor-based GPU cloud Lium buys back and burns nearly $2.7M of SN51 tokens in six months — markjeffrey · 2026-10-09
- Google open-sources ML Drift: one GPU engine for GLES/OpenCL/Metal/WebGPU, cutting Shorts frame latency 40% — lmoroney · 2026-10-09