KV Cache Splicing Method for Frozen Gemma 4
MindPsychological140 · reddit · 2026-07-19
The author introduces a method to save verified knowledge as KV states and restore them when needed, achieving byte-level identical results compared to recomputation.
They report that on Gemma 4 12B, this cached knowledge approach boosted a routing system's performance on AIME 2025 from 76.7% to 90.0%. The post includes a link to the paper (arXiv:2607.14431) and notes that the author will present at the AGI Summit on July 19.
Related event: KV-Cache Grafting Boosts Frozen Small Models(2 posts)→
More from Research
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- VidMap uses RoMa coarse matching on all frames, fine-scale only for keyframes — ducha_aiki · 2026-09-11
- Bug Hunt Bench author: leaderboard noise is about 2-3 points — PawelHuryn · 2026-09-11
- PNAS paper shows a tiny billiard-ball system is a universal computer — undecidability lives in two dimensions — eigensteve · 2026-09-11
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11