Tiered KV-cache offloading for self-hosted LLM inference: GPU to RAM to NVMe to S3
Responsible-You9024 · reddit · 2026-09-08
A practitioner is working on KV-cache offloading for self-hosted LLM inference: moving colder KV blocks down a tiered path — GPU → RAM → NVMe → S3 — instead of keeping everything in expensive VRAM. They're asking the community about real production experience: is the biggest pain VRAM capacity, latency when loading KV back, network bandwidth, or something else? Relevant reading for anyone optimizing local or large-scale inference.
More from Infra
- Actian's new vector DB claims 22x Qdrant speedup, but its docs look suspiciously like Qdrant's — qdrant_engine · 2026-09-08
- 8GB VRAM runs Flux and Wan 2.2 locally fine: three wrong settings, not the GPU, were the bottleneck — leonbuilds · 2026-09-08
- China effectively leads the humanoid robot supply chain, and Optimus relies on it — JOBhakdi · 2026-09-08
- Benchmarked: Apple Core AI vs MLX for on-device LLM speed on iPhone and Mac — HankYeomans · 2026-09-08
- D-Wave Finalizes Agreement with US Commerce Dept for Up to $100M in CHIPS Act Funding — ceciletamura · 2026-09-08
- Export Controls Working? H200 Sells for 280 and B300 for 450 Overseas — teortaxesTex · 2026-09-08