Increase llama.cpp -cram Beyond 8192MB for Agentic Workflows

My_Unbiased_Opinion · reddit · 2026-09-24

A Reddit PSA: llama.cpp's -cram prompt-cache default (8192MB) is too small for long multiturn agentic workflows, forcing full context reprocessing. The author found 20480 works well with Qwen 27B 3.8 at 262K context; the tradeoff is more RAM, not more VRAM.

Original post →

More from coding & agent

coding & agent channel →