KV cache transplants on Qwen3.8-27B: start at Q6, hand off to Q3, beat static quant
wadeAlexC · reddit · 2026-09-26
Inspired by the Cache-to-Cache paper (arXiv:2510.03215), which fuses KV caches between models via a trained converter, the author tested whether KV caches transfer across quantizations of the same model with no converter. On Qwen3.8-27B with NIAH-style tasks, dynamic strategies (start Q6K, swap to Q4KXL, then IQ3S, optionally quantizing KV cache to q80) fit in 24GB and beat running low-precision quants from the start on his benchmarks.
More from Infra
- How Long Until Local ~30B A3B Models Match GLM 5.3 Flash Quality? — Aggravating-Push-207 · 2026-09-26
- Vpipe vs Draw Things on M5 Pro: 24% faster at 1K, finishes 2K where Draw Things crashes — TgoAI · 2026-09-26
- AMD publishes educational GEMM optimization ladder for Helios MI455X GPUs with HipKittens — salykova_ · 2026-09-26
- Terafab starts hiring: 1 TW/year chip output and orbital AI compute in its sights — seanmcdonaldxyz · 2026-09-26
- Blog: Scaling LLM Inference from a Single Node to Millions — abhijithneil · 2026-09-26
- Samsung, Oxford and PKU propose TrOPD to distill frontier-model reasoning into on-device small models — jiqizhixin · 2026-09-26