Bolting on a small approximator halves Qwen3-8B prefill without retraining
rickasaurus · x · 2026-09-14
Borrowing DeepSeek V4.1 Flash's Causal Encoder-Decoder idea, researchers trained an external approximator module for Qwen3-8B that predicts the second-half KV cache, cutting prefill time nearly in half with identical outputs — no weight changes needed. If it generalizes, existing fine-tuned open models could get drastically faster local long-context deployment without waiting for next-gen releases.
More from Infra
- Qwen3.8-27B NVFP4 quants compared: lm_head precision makes or breaks MTP speedups — danielhanchen · 2026-09-14
- Musk says SpaceX will launch Nvidia AI computers into space next year — inductionheads · 2026-09-14
- Procedurally Generated FlashAttention With CuTe Tilings Yields Runnable CUDA Code — vtabbott_ · 2026-09-14
- Tahuna open-sources AI training infra for small teams: train, serve, orchestrate GPUs — Monaim101 · 2026-09-14
- Meta's MTIA 400 chip splits duties between LLM training and ad recommender inference — jonathanmendez · 2026-09-14
- Compute Middlemen: 'Uber Doesn't Have Cars Either' as a Playbook for Contracted GPU Capacity — MatthewChang · 2026-09-14