Apple's KV-Lingo translates KV caches between LLMs, cutting switch time up to 29x
Gauri_the_great · x · 2026-10-02
Switching LLMs mid-conversation normally requires recomputing the prefix KV cache, since caches are tied to each model's architecture. Apple's new KV-Lingo paper trains a linear translator that maps a source model's KV cache directly into a target model's representation space, bypassing re-prefill entirely.
How it works
- For each target layer, separate linear maps transform the source model's keys and values token-by-token; with mismatched depths, each target layer reads a selected subset of source layers.
- Training has two stages: a closed-form least-squares fit to the target's native cache, followed by self-distillation through the target model that optimizes next-token distribution agreement (KL divergence) rather than entry-wise cache reconstruction. Only the translator is trained; both models stay frozen.
- Translation cost scales linearly with prefix length and involves no attention computation.
Results
- On an H100 with shape-preserving Qwen3 pairs, switching is 4.8–5.6x faster than re-prefill at 2k tokens and 12–29x at 32k, with subsequent decoding at native speed.
- Across repeated switches, each model keeps its cache and only newly generated segments are translated.
- Across nine benchmarks (reasoning off), distilled linear translators roughly match the smaller model's average performance; for Qwen3-0.6B↔8B, head-mixing KV-Lingo scores 0.577 small-to-large and 0.476 large-to-small vs. the smaller model's 0.506.
The authors note quality remains dependent on the model pair and task — translated caches do not consistently recover the larger model's native performance.
More from Infra
- TensorFold creator seeks funding to go full-time on local AI performance work — EAccelerate_42 · 2026-10-02
- Building a permissioned, peer-to-peer AI inference network with Cascadia — techne98 · 2026-10-02
- Amazon pledges $1B in community support to win local backing for AI data centers — pstAsiatech · 2026-10-02
- Kirin 9050 die shot revealed: 120.6mm² on TSMC N2P, alongside 2026 flagships — zephyr_z9 · 2026-10-02
- LTX 2.5 generates 60-second clips on a 32GB RTX 5090 as software optimization beats VRAM upgrades — OpenEffect3955 · 2026-10-02
- AI agents that watch your screen hint at an eye-tracking, voice-driven future — perilli · 2026-10-02