KV-Cache Grafting Boosts Frozen Small Models
A new KV-cache grafting method stores verified knowledge as byte-exact KV states and restores them later, aiming to boost small frozen models without changing weights. The authors say the approach improves capability while cutting inference cost, with results reported on Gemma 4 12B.
2026-07-17 ~ 2026-07-19 · 2 related posts
- KV-cache Grafting: Freezing Small Models Makes Them Stronger — Corbenci · 2026-07-17
- KV Cache Splicing Method for Frozen Gemma 4 — MindPsychological140 · 2026-07-19