KV-cache Grafting: Freezing Small Models Makes Them Stronger

Corbenci · hf · 2026-07-17

Making Frozen Small Models Cheaper and Stronger with KV-cache Grafting

This work proposes a method to boost small model capabilities and reduce inference costs without changing weights: storing verified knowledge as byte-exact KV states, which are later grafted back into new inference contexts.

Core Features

Results

The authors describe this as a "verified-knowledge flywheel": knowledge is verified once and reused multiple times.

Related event: KV-Cache Grafting Boosts Frozen Small Models(2 posts)→

Original post →

More from Infra

Infra channel →