SemiAnalysis probes DeepSeek V4.1 Flash's Engram gates, revealing how the model activates text patterns
teortaxesTex · x · 2026-09-26
SemiAnalysis probed the Engram gates of DeepSeek V4.1 Flash to see which text patterns the model activates. The findings go well beyond names and facts, offering a rare look into its internal memory mechanisms.
Reposter teortaxesTex adds that DeepSeek is 'just that good,' but notes they probably still don't have 50K Hopper GPUs.
More from Infra
- NVIDIA at $5.4 trillion is now worth more than the entire UK or French stock market — iamfakhrealam · 2026-09-26
- DeepSeek V4.1 Flash's Engram memory layer trades FFN compute for lookup tables, SemiAnalysis data suggests — teortaxesTex · 2026-09-26
- A Curated Paper List for Learning Distributed LLM Training and Inference — East-Muffin-6472 · 2026-09-26
- vLLM adds Elastic Expert Parallelism: grow/shrink MoE GPU pools under live traffic — PyTorch · 2026-09-26
- AI data center investors now favor real infrastructure over PowerPoint promises — TansuYegen · 2026-09-26
- Running Qwen 27B Q4_K_M on dual RTX 3060 with llama.cpp hits ~44-50 tok/s — jacek2023 · 2026-09-26