Suffix cache reuse for hybrid attention lifts edited-turn hit rate 29.9% to 58.1%, cutting prefix-reuse FLOPs to 7.14/Q
RulinShao · x · 2026-10-02
Rulin Shao shares implementation details of Suffix Cache Reuse for hybrid attention (interleaved full and linear attention layers) built on SGLang:
- Hit-rate decomposition: context editing in CLMs tanks cache hit rate from 73.9% to 29.9% on edited turns.
- Fix: instead of discarding unmatched suffix cache, reuse it — adding 28.2% more hit rate on edited turns and dropping total prefix-reuse FLOPs from 10.98/Q to 7.14/Q at matched performance.
- Takeaway: the cache space allows more fancy manipulations beyond token space; a readable blog post is coming.
The efficiency metric accounts for re-prefill cost of un-hit KV cache, so the "CLMs are cheaper" claim already factors in lower hit rates.
More from Infra
- Top 10% of firms capture 99.5% of model-serving spend, AI compute data shows — bendee983 · 2026-10-02
- Extropic founder teases 'Thermo RSI is coming' in cryptic thermodynamic computing hype post — beffjezos · 2026-10-02
- Microsoft Foundry puts GPT-6, Claude Opus 5.5 and Grok-4.6 all in one catalog with 1M contexts — mustafasuleyman · 2026-10-02
- Musk: "Orbital compute is gonna be a very big deal" as SpaceX eyes space-based energy for AI — DimaZeniuk · 2026-10-02
- DeepSeek's DSec paper: 3M sandboxes/day, 380K concurrent, powering V3.2-V4.1 RL — jiqizhixin · 2026-10-02
- Greylock co-leads Series A in Parallax, an AI energy startup building 3D-printed gas turbines — SethGRosenberg · 2026-10-02