Rumored DeepSeek V4.1 Flash details point to asymmetric-activation MoE with much lower cost
gaganghotra_ · x · 2026-09-10
A viral repost claims unverified details about DeepSeek V4.1 Flash: a 552B MoE with a novel Causal-Encoder-Decoder architecture, asymmetric activation (8B reading, 16B writing) for much lower cost, plus new pretraining and large-scale RL posttraining allegedly beating V4 Pro and other flagships. Note these figures conflict with official/vLLM specs (769B total / 15.5B active) — treat with caution.
Related event: Alleged DeepSeek V4.1 and V4.1 Flash Benchmarks and Architecture Leak(12 posts)→
More from Models
- New Book Teaches Beginners to Build and Fine-Tune Their Own GPT-Style SLMs, With Colab Notebooks — Roger_M_Taylor · 2026-09-10
- DeepSeek tipped customers about V4.1 Flash ahead of open-weights launch — cedric_chee · 2026-09-10
- Apodex 1.1 mini lands on Hugging Face in GGUF for local deployment — SimonShaoleiDu · 2026-09-10
- Preliminary Assessment: Zhipu GLM V4.1 Envs Match V4, Post-Training More Advanced — xeophon · 2026-09-10
- Astra's DeepSeek V4 forecast vs what actually happened — teortaxesTex · 2026-09-10
- Leaked brief: DeepSeek V4.1 Flash at 552B params, foldable iPhone at $1,999, ChatGPT voice limits raised — testingcatalog · 2026-09-10