NVIDIA researchers show KV caches transfer between models via closed-form linear mapping

anselm · x · 2026-09-24

NVIDIA researchers found a way to transfer KV caches directly between models in the same family (e.g., Qwen3 14B → 32B), skipping the costly prefill when swapping from a small model to a larger one mid-conversation.

Original post →

More from Infra

Infra channel →