DeepSeek V4.1 Flash rumored to ship 196B-entry 'Engram' lookup table instead of computing everything
cephaloform · x · 2026-09-11
Per plotarmordev (unconfirmed), DeepSeek V4.1 Flash ships with a roughly 200GB lookup table called Engram: 196B parameters of stored entries sitting next to the 552B main model.
Instead of calculating everything, the model turns its last few tokens into an address and pulls matching rows like checking an index; each token reads only a few rows, so the table doesn't need expensive GPU memory. A reposter remarks it feels like alien technology after years of incremental 'transformer++' changes since Llama.
Related event: Leak: DeepSeek V4.1 Flash Is Actually a 718B Sparse Model(2 posts)→
More from Models
- DeepSeek Flash impresses: non-sycophantic, argumentative, and blazing fast — oran_ge · 2026-09-11
- Early take: DeepSeek V4.1 Solo beats Agent Teams and GLM 5.3 Flash on quality and cost — teortaxesTex · 2026-09-11
- What counts as an 'exchange'? Anthropic's 865K/day metric questioned — teortaxesTex · 2026-09-11
- OpenAI reportedly aims its new internal model at Riemann and P vs NP — zephyr_z9 · 2026-09-11
- MLX vs CUDA: Qwen 3.8 Flash Next optimization duel delivers 55%+ speedups on both sides — gajesh · 2026-09-11
- GPT-6 Astra drives a robot arm on first try via physical ICL; Ken Goldberg touts Agentic Robotics — zhaoran_wang · 2026-09-11