DeepSeek 4.1 flash reportedly uses large ngram embeddings, echoing Qwen4 architecture

ccerrato147 · x · 2026-09-11

nisten notes that the new DeepSeek 4.1 flash also uses very large ngram embeddings (storable in CPU RAM), similar to the upcoming Qwen4 architecture, and shares a full 3D breakdown of every weight in the model and what it does. Unconfirmed third-party observation.

Original post →

More from Models

Models channel →