llama.cpp reasoning-preserve flag risks context bloat
anderspitman · reddit · 2026-08-19
A Reddit user discussed the --reasoning-preserve flag in llama.cpp.
This setting causes llama.cpp to include full reasoning traces in the conversation history instead of just the answers. While this sounds like it would improve quality, the user is concerned that the sheer volume of the model's thoughts might significantly increase context usage.
More from Infra
- NVIDIA publishes multi-GPU method for massive-scale UMAP in minutes — leland_mcinnes · 2026-08-19
- Local inference economics: $60/month power bill for slow speeds — Thin_Pollution8843 · 2026-08-19
- Australia offers free midday power, challenging space datacenter economics — aronchick · 2026-08-19
- DFlash2 on Qwen3.8 27B hits ~200tk/s for code, requires more VRAM — Hefty_Wolverine_553 · 2026-08-19
- ZML runtime now supports 9 hardware platforms including NVIDIA, AMD, and MooreThreads — ylecun · 2026-08-19
- Cerebras holds first conference as public co, chips power OpenAI — Scobleizer · 2026-08-19