DeepSeek update keeps cache hits mid-conversation, cuts costs 36.6%
teortaxesTex · x · 2026-09-11
- A hands-on test shows DeepSeek V4.1 Flash with the new DSH harness no longer busts prefix cache when the System Prompt changes mid-conversation.
- Across 20 six-turn comparisons in 5 scenarios: total tokens up only 2.4% while cost fell 36.6%; cache hit rate jumped from 0% to 88.3%, non-cached input dropped 56.8%.
- Author notes Codex already had similar behavior; DeepSeek is catching up on both model and harness.
More from Models
- Qwen3-8B gets a KV-approximation add-on that halves prefill time without touching the model — teortaxesTex · 2026-09-11
- Pro 20x tier burns 60% of weekly quota in under a day with GPT-6 Astra — rschu · 2026-09-11
- Is DeepSeek's rumored K3 a scaled-down model, or something bigger? X users debate — teortaxesTex · 2026-09-11
- 6TB of Fable data sold with leaked SSH keys, cloud creds tied to Xiaomi, Huawei, NIO — teortaxesTex · 2026-09-11
- AI Sextet offers 6 models free and unlimited for 14 days, including DeepSeek and Qwen — airesearch12 · 2026-09-11
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11