llama.cpp reasoning-preserve flag risks context bloat

anderspitman · reddit · 2026-08-19

A Reddit user discussed the --reasoning-preserve flag in llama.cpp.

This setting causes llama.cpp to include full reasoning traces in the conversation history instead of just the answers. While this sounds like it would improve quality, the user is concerned that the sheer volume of the model's thoughts might significantly increase context usage.

Original post →

More from Infra

Infra channel →