99.7% cache hits: engineered DeepSeek Harness with self-hosted GLM-5.3

burny_tech · x · 2026-09-21

Ahmad Osman reports getting 99.7% cache hits running a self-hosted GLM-5.3 with DeepSeek Harness, crediting the result to quality-engineered harnesses rather than vibe-coded ones — context layout directly determines prompt cache hit rates and inference cost.

Original post →

More from coding & agent

coding & agent channel →