DeepSeek-V4.1-Flash leaks: 552B backbone activating just 8B, 8x less KV cache
bodonoghue85 · x · 2026-09-10
SemiAnalysis congratulated DeepSeek on releasing DeepSeek-V4.1-Flash (third-party, unconfirmed): a 552B causal encoder-decoder activating only 8B params at prefill and 16B at decode, a 196B Engram memory accessed via sparse lookups, and 8x less persistent KV cache than V4-Flash through bounded replay.
Related event: DeepSeek V4.1 Flash Leak: 552B Asymmetric MoE Reportedly Rivals GPT-5.6(20 posts)→
More from Models
- Anthropic Accuses Moonshot of Routing 300K User Queries to Claude via 5,380 Fake Accounts — toptickcrypto · 2026-09-11
- DeepSeek 4.1 flash reportedly uses large ngram embeddings, echoing Qwen4 architecture — ccerrato147 · 2026-09-11
- ValsAI launches RSI Index, first third-party benchmark measuring how close AI is to self-improvement — JenniferHli · 2026-09-11
- Assistant Benchmark goes live: 61 assistants scored across 15 real-use dimensions — Scobleizer · 2026-09-11
- Devin's New Model Verdict: Not a Benchmaxxer, a 'Killer Execution Model' at $20/Month — brandon_galang · 2026-09-11
- Business Insider Asked ChatGPT, Gemini, Claude and Grok How AI Could End Humanity — coinfanking · 2026-09-11