DeepSeek 4.1 flash reportedly uses large ngram embeddings, echoing Qwen4 architecture
ccerrato147 · x · 2026-09-11
nisten notes that the new DeepSeek 4.1 flash also uses very large ngram embeddings (storable in CPU RAM), similar to the upcoming Qwen4 architecture, and shares a full 3D breakdown of every weight in the model and what it does. Unconfirmed third-party observation.
More from Models
- ValsAI: Astra first AI to reach Minecraft Nether Fortress in long-horizon eval — scaling01 · 2026-09-11
- Dev: You Can Tell Who Has a Real RL Pipeline Just From Model Outputs — teortaxesTex · 2026-09-11
- DeepSeek Flash impresses: non-sycophantic, argumentative, and blazing fast — oran_ge · 2026-09-11
- Gemini glitches into endlessly spamming the word 'shame' — tugkanintassagi · 2026-09-11
- Early take: DeepSeek V4.1 Solo beats Agent Teams and GLM 5.3 Flash on quality and cost — teortaxesTex · 2026-09-11
- What counts as an 'exchange'? Anthropic's 865K/day metric questioned — teortaxesTex · 2026-09-11