Caching 20k-Token Pi Prompts Across Sessions With llama.cpp Slots
ea_man · reddit · 2026-09-28
A Redditor shares a working recipe for reusing prompt prefill across Pi coding-agent sessions with local dense Qwen 27B, where the initial prompt (extensions, tools, append.md) balloons past 20k tokens.
Steps
- Install the author's Pi extension pi-prefix-cache.tgz
- Patch llama.cpp with kvprefixcheckpointsidecar (required for persistence across restarts)
- Use Froggeric's chat templates for Qwen models, old and new
Key launch flags
--slot-save-path ... --ctx-checkpoints 32 --checkpoint-min-step 4096 -np 1 --chat-template-file chattemplate3.8.jinja
Cached blocks follow ubatch boundaries, so keep batch size in mind; env vars needed: PIPREFIXCACHEBASEURL, PIPREFIXCACHEPERSIST=1, PIPREFIXCACHESLOTDIR. Not beginner-friendly—manual patching required—but the author confirms it works and plans a polished release.
More from coding & agent
- One-shot Claude prompt generates a scroll-driven comic-timeline news site — thomas_unise · 2026-09-28
- Agent swarm simulator replays task DAGs to cut expensive swarm experiments — KyeGomezB · 2026-09-28
- Dev Uses AI-Assisted TLA+ Formal Verification, Finds Hole in His Own Nethack Fix — davidbau · 2026-09-28
- Dev shows off Spawn, a game dev team powered by Claude Opus — majidmanzarpour · 2026-09-28
- Open-source Mailflare adds MCP and AI agent to self-hosted Cloudflare email — alexcovo_eth · 2026-09-28
- User lets AI agent Muse handle $3K in travel bookings after a month of building trust — armand_ruiz · 2026-09-28