Suffix cache reuse for hybrid attention lifts edited-turn hit rate 29.9% to 58.1%, cutting prefix-reuse FLOPs to 7.14/Q

RulinShao · x · 2026-10-02

Rulin Shao shares implementation details of Suffix Cache Reuse for hybrid attention (interleaved full and linear attention layers) built on SGLang:

The efficiency metric accounts for re-prefill cost of un-hit KV cache, so the "CLMs are cheaper" claim already factors in lower hit rates.

Original post →

More from Infra

Infra channel →