DeepSeek V4.1 Flash is actually 748B params, safetensors analysis shows
DistanceSolar1449 · reddit · 2026-09-10
A Reddit user inspected the safetensors on Hugging Face to settle confusion over DeepSeek V4.1 Flash's size: the main model is 551.566B params across 40 layers, with FFN experts totaling 543.582B and the rest (attention, shared experts, etc.) just 7.984B.
HuggingFace lists 485B because it counts some FP4-packed weights as bytes rather than params (2 FP4 params per byte). On top: the engram is 196.929B, DSpark/MTP 14.225B, and the vision encoder only 0.485B (much smaller than expected) — technically optional. All-in it's 748B, so 128GB or 256GB of RAM/VRAM won't cut it.
More from Models
- Analyst: DeepSeek's latest change is a big win for token efficiency, moving toward OpenAI's regime — teortaxesTex · 2026-09-10
- antirez: stellar cybersec benchmarks aside, wait for real user testing before trusting the best scores — antirez · 2026-09-10
- antirez on the new DeepSeek model: 2-bit quants may not hold up, and it's not really 'local' — antirez · 2026-09-10
- DeepSeek report: post-training gains come from better data and environments, not RL novelty — realsohamparekh · 2026-09-10
- 'The whale is back': DeepSeek reportedly releases a new report — scaling01 · 2026-09-10
- Engram architecture explained: 500B backbone beats GLM 5.3-class with far smaller KV cache — bookwormengr · 2026-09-10