DeepSeek-V4-Flash Tip: Mid-conversation System Messages Break Prompt Cache
CharlesStross · reddit · 2026-08-02
Developer Charles Stross shared an important PSA for DeepSeek-V4-Flash users: inserting system role messages mid-conversation will break your prefix cache, degrading performance and increasing costs.
The Cause: DeepSeek's official chat template format lacks a mid-conversation system turn. Forcing a system message in the middle or at the tail of a conversation disrupts the context structure, effectively frying the prefix cache.
The Solution: Users should utilize the latestreminder role instead of system for mid-conversation instructions. This is the role DeepSeek trained for this purpose, and it passes through inference engines like llama.cpp without issue, maintaining cache hit rates.
More from Models
- Tencent Releases UI-Mate-27B, a Desktop GUI Agent Model — tencent · 2026-08-24
- Sakana AI translation outperforms Google and DeepL in Japanese-English benchmarks — SakanaAILabs · 2026-08-24
- Developer haider makes his own LLM tier list after disagreeing with theo's rankings — haider1 · 2026-08-24
- Mystery OxAlpha Beats Claude; Alibaba Raises $10B for AI — 创业邦 · 2026-08-24
- OpenAI and Google cut LLM prices; mystery OxAlpha model beats Claude on DeepSWE — 创业邦 · 2026-08-24
- AI News Digest: DeepSeek Weekend Discounts, GPT-5.6 Sol Price Cut, Alibaba's $10B AI Raise — APPSO · 2026-08-24