Persisting KV Cache on Free ARM: 15x Faster LLM Prefill

Annual_Manner_5901 · reddit · 2026-08-03

Testing on a free Oracle ARM box (4 cores), the author found that CPU inference is bottlenecked by the prefill phase. By persisting the KV cache to disk, a new process can inherit the prefill, slashing time from 54.4s to 3.5s (or 0.1s from page cache) for 3356 tokens.

Bugs & Failed Attempts

Surprising Wins

Original post →

More from Infra

Infra channel →