Apple's KV-Lingo translates KV caches between LLMs, cutting switch time up to 29x

Gauri_the_great · x · 2026-10-02

Switching LLMs mid-conversation normally requires recomputing the prefix KV cache, since caches are tied to each model's architecture. Apple's new KV-Lingo paper trains a linear translator that maps a source model's KV cache directly into a target model's representation space, bypassing re-prefill entirely.

How it works

Results

The authors note quality remains dependent on the model pair and task — translated caches do not consistently recover the larger model's native performance.

Original post →

More from Infra

Infra channel →