Test: DeepSeek Harness achieves 99% cache hit rate with GLM and Kimi
sandyyevans · reddit · 2026-08-19
An experiment tested DeepSeek Harness (DSH) with GLM, Kimi, and Opus to verify if prompt caching works with non-DeepSeek models. Results showed GLM achieved 97% cache reuse in tool loops and 99.6% in subsequent turns, while Kimi hit 99% in both. This confirms DSH's append-style request pattern is compatible with other providers supporting prefix caching. Opus showed no cache activity in this test.
More from Infra
- Dyna-2 trained on 1M+ hours of egocentric video; data infra detailed, open-sourced — Scobleizer · 2026-08-19
- NVIDIA open-sources TensorRT Model Connect, two commands to TensorRT inference — JFPuget · 2026-08-19
- Local LLM Noise: Running Bots at Home Sounds Like a Jet Engine — Daniel_Farinax · 2026-08-19
- MiniMax H3 on a 16GB M5 MacBook Air: VPipe Beats h3.c by ~25% Wall-Clock — TgoAI · 2026-08-19
- NVIDIA monopoly hard to break? Paradigm shifts are key — 2C_ornot2C · 2026-08-19
- Anthropic's Multi-Level Monitoring for Astra Inference Revealed — AccBalanced · 2026-08-19