KV cache transplants on Qwen3.8-27B: start at Q6, hand off to Q3, beat static quant

wadeAlexC · reddit · 2026-09-26

Inspired by the Cache-to-Cache paper (arXiv:2510.03215), which fuses KV caches between models via a trained converter, the author tested whether KV caches transfer across quantizations of the same model with no converter. On Qwen3.8-27B with NIAH-style tasks, dynamic strategies (start Q6K, swap to Q4KXL, then IQ3S, optionally quantizing KV cache to q80) fit in 24GB and beat running low-precision quants from the start on his benchmarks.

Original post →

More from Infra

Infra channel →