Llama.cpp restarts kill long context prefills: proposal for auto KV state persistence
Dazzling_Equipment_9 · reddit · 2026-08-20
Running local models on a 128GB Strix Halo, the user faces slow 50k-100k context refills after restarting llama.cpp during testing. While llama-server offers slot save/restore APIs, they suggest a server-side mechanism to automatically cache KV/prefill states to disk and restore them upon session return. This generic cache (evicted by LRU) would significantly improve UX for slower-prefill hardware without requiring client-side integration.
More from Infra
- 4th generation cryostat unveiled with 250+ photonic chips — jwt0625 · 2026-08-20
- Ethernet copper costs as low as $0.10-0.25/Gbps — jwt0625 · 2026-08-20
- Nvidia's Sustained Performance Limited by Power Delivery: 9.7 PFLOPS — nabla_theta · 2026-08-20
- ML Researcher: Compute Shortage, Not Ideas, Slowing Down Progress — A_K_Nain · 2026-08-20
- M3 Ultra achieves 45.8% performance gain running 27B model with CPU+GPU+ANE offload — AIFlow_ML · 2026-08-20
- YC-backed Atomarine builds floating data centers to solve the compute crunch — ycombinator · 2026-08-20