Increase llama.cpp -cram Beyond 8192MB for Agentic Workflows
My_Unbiased_Opinion · reddit · 2026-09-24
A Reddit PSA: llama.cpp's -cram prompt-cache default (8192MB) is too small for long multiturn agentic workflows, forcing full context reprocessing. The author found 20480 works well with Qwen 27B 3.8 at 262K context; the tradeoff is more RAM, not more VRAM.
More from coding & agent
- Meta's Proactive Memory Agent fixes context rot, lifting Claude Sonnet 4.5 from 37.6% to 45.9% — DeepLearningAI · 2026-09-25
- A practical guide to Claude Code 2.0: sub-agents, hooks, and context engineering — dejavucoder · 2026-09-25
- mitsuhiko: The Experience of Models Editing Files via Code Is Still Unacceptable — mitsuhiko · 2026-09-25
- Nimble launches official OpenAI plugin for ChatGPT and Codex — OpenAIDevs · 2026-09-25
- Asupersync: A Cancel-Correct Async Runtime for Rust, Built Heavily with Frontier Models — doodlestein · 2026-09-25
- Building Macross Plus Mecha with Codex: A Dev's Side Project — algo_diver · 2026-09-25