BeaconKV compresses KV cache for long reasoning models via beacon queries
Janghyeon Kim · hf · 2026-09-09
BeaconKV uses compact beacon queries to predict which past key-value pairs will be revisited during long reasoning traces, shrinking KV cache size without sacrificing accuracy—an inference-efficiency method for large reasoning models.
More from Infra
- One flag lets vLLM serve any Transformers model — no handwritten port needed — ariG23498 · 2026-09-09
- OpenAI may pause new Pro subscriptions as demand for Astra hits unprecedented levels — op7418 · 2026-09-09
- The 4 things you need to run local AI: models, Hugging Face, runners, and quantization — Roger_M_Taylor · 2026-09-09
- Google says AI servers pay back in under 2 years, just 1 year on its own silicon — SumitGup · 2026-09-09
- Together claims GLM-5.3 Flash beats Claude Fable 5.1 on agentic tasks at ~1% cost — togethercompute · 2026-09-09
- A Biography of Lee Holloway, the Architect of Cloudflare's Technology (Part 1) — porridgeraisin · 2026-09-09