PReCache: Training-Free KV Cache Sharing Gives Multi-LoRA Agents up to 3.1x TTFT Speedup
SNU-VLSI · hf · 2026-09-29
Multi-LoRA agent systems specialize roles on a shared backbone, but each agent reprocesses the growing shared trajectory and builds its own KV cache—heavy redundancy in long-horizon tasks, and direct reuse weakens the current agent's LoRA-specific behavior. SNU-VLSI's training-free PReCache adds two designs:
- PreLRShared: precomputes a compact per-agent low-rank cache when shared context is first processed with pretrained weights, letting each agent use its own LR cache without repeated prefill
- ReBaseShared: reconstructs the shared base cache from adapter-free hidden states to cut error from the previous agent's adapted representations, with two inference schemes for single-stream and concurrent serving
- Up to 3.1x TTFT speedup and 2.3x per-request throughput vs. no cache sharing; ReBaseShared preserves accuracy best, dropping only 1.1 points on average
More from Infra
- Qdrant unveils Constella research preview: swap query embedding models without re-embedding your docs — qdrant_engine · 2026-09-29
- 124M model with a 65B embedding sparks the AFED disaggregation joke — YouJiacheng · 2026-09-29
- Oracle's 30-year spread widens to +180bp as Project Jupiter power woes trigger force majeure — julsimon · 2026-09-29
- Bain: AI needs $6T annual revenue by 2031 to justify data-center spending, $4.2T gap remains — rohanpaul_ai · 2026-09-29
- Cost math: Meta Muse would run $51 per user, making 'free for 4B users' a steep climb — bookwormengr · 2026-09-29
- KVCMAS Corrects Shared-Context KV Cache Online, 2.0x TTFT Speedup for Multi-Agent Serving — SNU-VLSI · 2026-09-29