Microsoft's OasisKV boosts LLM inference throughput 1.69x with lookahead sparse prefetching
microsoft · hf · 2026-08-11
Microsoft introduces OasisKV, a memory-centric LLM inference system that exploits sparse attention to keep only relevant KV entries in HBM, prefetching future important tokens via speculative decoding. Built on vLLM, it maintains accuracy within 0.7 points under a 2048-token KV budget, achieves 1.69x throughput over dense vLLM on reasoning workloads (0.1 point accuracy loss), up to 2.1x on multi-GPU long-context serving, and about 2x throughput with 6.5-9.7x less KV admission under prefill-decode disaggregation.
More from Infra
- Muse Glimmer 30B Tested at 1M Context: Perfect Retrieval on Consumer Hardware — StartupTim · 2026-08-11
- vLLM Muse Glimmer speculative decoding needs 6 patches, boosts speed from 25 to 57 tok/s — j4ys0nj · 2026-08-11
- JPMorgan Investors Expect FY27 HBM Contract Pricing to Jump Over 50% — zephyr_z9 · 2026-08-11
- 4x DGX Spark Cluster Achieves 44.6 tok/s on GLM-5.2 for Real Agent Workloads — EAccelerate_42 · 2026-08-11
- NVIDIA AI Hardware Fabric Roadmap: Transitioning from Glass to Pure Quartz — zephyr_z9 · 2026-08-11
- Ditching the Cloud: A Local Multi-Agent Setup with Large and Small Models — gnukeith · 2026-08-11