DeepSeek v4.1-Flash: 763B causal encoder-decoder with vision, KV cache cut to 1/8

Latent Space · rss · 2026-09-12

Latent Space's deep dive on DeepSeek v4.1-Flash: a 763B model with a novel causal encoder-decoder architecture, splitting prefill (8B active) from decode (16B active) for 1-2% sparsity. With Sliding-Window Attention Bounded Replay, KV cache footprint drops to as little as 1/8 of V4 Flash, making it faster and cheaper for long-running agents.

Key points:

Original post →

More from Models

Models channel →