DeepSeek Releases V4.1-Flash with 1M Token Context
DeepSeek released V4.1-Flash, its smallest new-architecture model, featuring a 1M token context, 4x smaller KV cache, and a multimodal MoE design with 552B backbone parameters (763B total). It is now available on the FLock API platform and Ollama's cloud mode.
2026-09-11 ~ 2026-09-11 · 2 related posts
- Episode 1: DeepSeek V4.1 Flash Opens Limited Internal Beta with New Architecture and Native Multimodality(2026-09-08, 18 posts)
- Episode 2: DeepSeek V4.1 Flash Tested: Blazing 350 Tokens/s but Still Experimental(2026-09-08, 2 posts)
- Episode 3: DeepSeek cuts V4-Flash API prices with new peak/off-peak billing from Sept 10(2026-09-08, 6 posts)
- Episode 4: DeepSeek V4.1 Flash Nears Frontier Models at a Fraction of the Cost, Independent Tests Show(2026-09-09, 15 posts)
- Episode 5: DeepSeek Releases Open-Source V4.1 Flash, Beating Its Own Flagship at a Fraction of the Cost(2026-09-09, 57 posts)
- Episode 6: DeepSeek V4.1 Flash Leak: 552B Asymmetric MoE Reportedly Rivals GPT-5.6(2026-09-10, 20 posts)
- Episode 7: DeepSeek V4.1 Flash Benchmark Results Leak, Sparking Community Discussion(2026-09-10, 2 posts)
- Episode 8: DeepSeek Compresses KV Cache 54x in Nine Months to Sub-KB Per Token(2026-09-10, 6 posts)
- Episode 9: Leaked DeepSeek-V4.1-Flash Touted as New Open-Source King(2026-09-10, 3 posts)
- Episode 10: Bug Hunt Bench Tests Frontier Models on 105 Real Planted Bugs(2026-09-10, 8 posts)
- Episode 11: DeepSeek V4.1 Flash Confirmed to Center on YOCO Architecture(2026-09-10, 7 posts)
- Episode 12: DeepSeek V4.1 Flash Tech Report Leaks: KV Cache Compression Obsession Across the Stack(2026-09-11, 42 posts)
- Episode 13: DeepSeek V4.1 Flash tops Vals open-source index at $0.30 per run(2026-09-11, 2 posts)
- Episode 14: DeepSeek V4.1 Flash scores 40 in AA index despite rising hallucinations(2026-09-11, 2 posts)
- Episode 15: Inside KV-Cache Sharing: How DeepSeek CED and GLM 5.2 Differ(2026-09-11, 4 posts)
- Episode 16: DeepSeek Releases V4.1-Flash with 1M Token Context(2026-09-11, 2 posts)
- DeepSeek-V4.1-Flash hits Ollama: 552B MoE backbone with 1M context via KV cache compression — ollama · 2026-09-11
- DeepSeek launches V4.1-Flash with 1M-token context and 4x smaller KV-cache — matlabulous · 2026-09-11