FULL STORY
DeepSeek V4.1 Flash: Launch, Benchmarks and Cheaper Inference
DeepSeek V4.1 Flash launched on Together AI, claiming to beat GPT-5.6 Sol at a third of the cost. Independent benchmarks on 8×A40 followed, and DeepSeek then introduced a KV cache compression method to further cut inference memory costs.
2026-09-12 ~ 2026-09-13 · 3 episodes · 7 posts
Episode 1 · DeepSeek V4.1 Flash Launches on Together AI, Beating GPT-5.6 Sol at a Third of the Cost (2026-09-12, 3 posts)
DeepSeek V4.1 Flash is now available on Together AI, claiming agentic benchmark wins over GPT-5.6 Sol and V4 Pro at one-third the per-task cost. Built on a 552B+196B MoE with 890B KV-cache compression, it reportedly beats a 1.6T-parameter model with about a third of the parameters.
- DeepSeek V4.1 Flash tech report: KV cache compression lets 552B beat 1.6T — 机器之心 · 2026-09-12
- DeepSeek V4.1 Flash lands on Together AI at one-third the cost per task — togethercompute · 2026-09-13
- DeepSeek V4.1 Flash: 890 bytes/token KV cache, 552B MoE beats V4 Pro on agentic tasks — togethercompute · 2026-09-13
Episode 2 · TensorSharp Runs DeepSeek V4.1 Flash on 8x A40 with Strong Results (2026-09-13, 2 posts)
The open-source TensorSharp engine added a DeepSeek V4.1 Flash execution path, achieving roughly 40 tok/s decode and over 500 tok/s prefill on 8x A40 GPUs with GGUF quantization.
- DeepSeek V4.1 Flash Hits 40 tok/s on 8x A40 with Open-Source TensorSharp Engine — fuzhongkai · 2026-09-13
- TensorSharp Hits 500+ tok/s Prefill for DeepSeek V4.1 Flash on 8×A40 GPUs — fuzhongkai · 2026-09-13
Episode 3 · DeepSeek V4.1-Flash Compresses KV Cache to 890 Bytes per Token (2026-09-13, 2 posts)
DeepSeek's V4.1-Flash introduces a new KV cache compression method that cuts memory to just 890 bytes per token, enabling longer contexts and significantly lower inference memory costs.
- DeepSeek V4.1 Flash cuts KV-cache to 890 bytes per token for cheap long context — Prompt Engineering · 2026-09-13
- DeepSeek's V4.1-Flash KV cache compression could undercut OpenAI and Anthropic's compute moat — justlikemedics · 2026-09-13