DeepSeek V4.1-Flash squeezes KV cache to 890 bytes per token, cuts persistent cache 8x
jbhuang0604 · x · 2026-10-01
- DeepSeek V4.1-Flash compresses global KV cache to just 890 bytes per token and shrinks the persistent cache 8x, while staying competitive on agentic benchmarks.
- Three key techniques enable this: Causal Encoder-Decoder, Compressed Sparse Attention, and SWA Bounded Replay.
- Significant implications for inference cost on long-context and agent workloads; the linked video walks through the details.
More from Infra
- Omarchy on WSL: patched aquamarine backend brings Hyprland to Windows with GPU rendering — sytelus · 2026-10-01
- Qwen-Image 2.1 prompt enhancer hits 4.4x speedup in ComfyUI, now runs on 8GB VRAM — mozophe · 2026-10-01
- Running Omarchy desktop in Windows via WSL with GPU acceleration and 4K multi-monitor support — sytelus · 2026-10-01
- Trader initiates Cerebras position, betting SRAM-based inference beats HBM as agents multiply model calls — Sethwinterroth · 2026-10-01
- Cerebras bull case: OpenAI paid tier, ~750 tok/s, $20B+ potential value and $25B RPO — Sethwinterroth · 2026-10-01
- MLX-Serve 26.10.1 ships with up to 66% faster Qwen3.8 27B inference on Apple Silicon — TheMoonMidas · 2026-10-01