99.7% cache hits: engineered DeepSeek Harness with self-hosted GLM-5.3
burny_tech · x · 2026-09-21
Ahmad Osman reports getting 99.7% cache hits running a self-hosted GLM-5.3 with DeepSeek Harness, crediting the result to quality-engineered harnesses rather than vibe-coded ones — context layout directly determines prompt cache hit rates and inference cost.
More from coding & agent
- Workflow tip: dictate rambling prompts to ChatGPT, then paste into your agents — athyuttamre · 2026-09-21
- Jesse agent beats Minecraft Ender Dragon in 7:06 on first try, cost $0.00 — BLUECOW009 · 2026-09-21
- Dev spends 3 days reverse-engineering Instinct's memory system, recreates it in 60 lines — iamrobotbear · 2026-09-21
- Code Contracts: open format puts eval-like specifications next to your code for AI agents — aronchick · 2026-09-21
- Codex users blast OpenAI over opaque usage resets: 'tell us the limit and refill rate' — StewartalsopIII · 2026-09-21
- Berating GPT vs. coddling a local model: two opposite prompting styles — gaviniboom · 2026-09-21