FULL STORY

DeepSeek V4.1 Flash: Launch, Benchmarks and Cheaper Inference

DeepSeek V4.1 Flash launched on Together AI, claiming to beat GPT-5.6 Sol at a third of the cost. Independent benchmarks on 8×A40 followed, and DeepSeek then introduced a KV cache compression method to further cut inference memory costs.

2026-09-12 ~ 2026-09-13 · 3 episodes · 7 posts

Episode 1 · DeepSeek V4.1 Flash Launches on Together AI, Beating GPT-5.6 Sol at a Third of the Cost (2026-09-12, 3 posts)

DeepSeek V4.1 Flash is now available on Together AI, claiming agentic benchmark wins over GPT-5.6 Sol and V4 Pro at one-third the per-task cost. Built on a 552B+196B MoE with 890B KV-cache compression, it reportedly beats a 1.6T-parameter model with about a third of the parameters.

Episode 2 · TensorSharp Runs DeepSeek V4.1 Flash on 8x A40 with Strong Results (2026-09-13, 2 posts)

The open-source TensorSharp engine added a DeepSeek V4.1 Flash execution path, achieving roughly 40 tok/s decode and over 500 tok/s prefill on 8x A40 GPUs with GGUF quantization.

Episode 3 · DeepSeek V4.1-Flash Compresses KV Cache to 890 Bytes per Token (2026-09-13, 2 posts)

DeepSeek's V4.1-Flash introduces a new KV cache compression method that cuts memory to just 890 bytes per token, enabling longer contexts and significantly lower inference memory costs.