Caching 20k-Token Pi Prompts Across Sessions With llama.cpp Slots

ea_man · reddit · 2026-09-28

A Redditor shares a working recipe for reusing prompt prefill across Pi coding-agent sessions with local dense Qwen 27B, where the initial prompt (extensions, tools, append.md) balloons past 20k tokens.

Steps

Key launch flags

--slot-save-path ... --ctx-checkpoints 32 --checkpoint-min-step 4096 -np 1 --chat-template-file chattemplate3.8.jinja

Cached blocks follow ubatch boundaries, so keep batch size in mind; env vars needed: PIPREFIXCACHEBASEURL, PIPREFIXCACHEPERSIST=1, PIPREFIXCACHESLOTDIR. Not beginner-friendly—manual patching required—but the author confirms it works and plans a polished release.

Original post →

More from coding & agent

coding & agent channel →