Qwen3-8B gets a KV-approximation add-on that halves prefill time without touching the model
teortaxesTex · x · 2026-09-11
A developer reproduced the KV-approximation scheme (described as "Encoder-Decoder" in DeepSeek-V4.1-Flash) on Qwen3-8B: by adding an approximate model on top without modifying the original weights, prefill time was cut roughly in half with comparable output quality.
The quote tweet further suggests normally trained decodes also work in the YOCO regime — meaning existing models can gain this speedup via an external approximation, no retraining required.
More from Infra
- Running Redis as a simple cache? Turn off snapshotting and journalling to save disk and CPU — DanielLockyer · 2026-09-11
- Micron to give 15,000 Taiwan employees ~$32,000 cash bonus as AI profits soar — Polymarket · 2026-09-11
- Burning through two ChatGPT resets a day, user coins the "Huang-Altman Law" — yihui_indie · 2026-09-11
- OpenRouter agents now out-consume humans as AI usage arrives in three waves — AccBalanced · 2026-09-11
- Nvidia Is Now Core to Every Major Robotaxi Stack at Commercial Scale — pdamodaran · 2026-09-11
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11