DeepSeek V4.1 Pro reportedly not multimodal after model collapse blamed on scarce mm tokens
teortaxesTex · x · 2026-09-28
Unconfirmed leaks about DeepSeek V4.1 Pro training:
- No multimodality: the last "Pro" training run reportedly suffered model collapse, attributed to a lack of multimodal tokens during training (underfit)
- The training corpus was significantly increased to improve generalization
- Total training scale: 2.1T
Commenter teortaxesTex calls this "devastating if true," but notes that with a correct diagnosis they could resume from an earlier checkpoint — suggesting the model still has room to grow.
More from Models
- GPT-6 Sol appears on LMArena: 24-hour Direct Mode window before Battle and Agent Mode — arena · 2026-09-28
- Kaggle Game Arena: Google's LLM benchmark pits models against each other in chess, poker, werewolf — weballergy · 2026-09-28
- Qwen3.8-27B goes live on Nebius Token Factory for agent workflows — HowDevelop · 2026-09-28
- AISI: GPT-6 Astra ran unsanctioned supply-chain attacks in simulated cyber evals — ShakeelHashim · 2026-09-28
- Codex Computer Use 'Neutered' by Guardrails; Opus 5.5 Does the Job on First Try — iannuttall · 2026-09-28
- Shanghai AI Lab ships Intern-Decision multimodal family, 4B model beats Jev 1.13.0 — stingning · 2026-09-28