KVCMAS Corrects Shared-Context KV Cache Online, 2.0x TTFT Speedup for Multi-Agent Serving
SNU-VLSI · hf · 2026-09-29
Prompt-specialized multi-agent systems share one model, but agent-specific prefixes alter the KV cache of the same shared context, forcing repeated prefills and duplicated caches. SNU-VLSI's KVCMAS is an online KV cache correction framework:
- Represents cross-agent cache deviations with compact low-rank states and chains corrections along the agent workflow, with no extra reference prefill
- Supports dynamically changing shared context while preserving an exact first-agent cache
- Matches or improves accuracy over prior sharing methods with the lowest TTFT under high concurrency
- 2.0x TTFT speedup over no cache sharing; up to 3.7x lower peak GPU memory than a prior correction method
More from Infra
- Hugging Face Transformers now runs llama.cpp GGUF quants natively on your laptop — ariG23498 · 2026-09-29
- Uber Eats ranking models serve 8M predictions/sec: how Uber scales ML feature consistency — AxSaucedo · 2026-09-29
- Qdrant unveils Constella research preview: swap query embedding models without re-embedding your docs — qdrant_engine · 2026-09-29
- 124M model with a 65B embedding sparks the AFED disaggregation joke — YouJiacheng · 2026-09-29
- Oracle's 30-year spread widens to +180bp as Project Jupiter power woes trigger force majeure — julsimon · 2026-09-29
- Bain: AI needs $6T annual revenue by 2031 to justify data-center spending, $4.2T gap remains — rohanpaul_ai · 2026-09-29