DeepSeek multimodal team praised for unusually careful pre-training data work
zephyr_z9 · x · 2026-09-10
Following DeepSeek's new model release, longtime China-AI observer teortaxesTex congratulated the DeepSeek multimodal team for finally delivering a truly premium model, saying criticism of their multimodal efforts should stop. Quoting the thread, eliebakouch highlighted the extreme care taken with pre-training data, something he says he hasn't seen in other models.
More from Models
- Leaked brief: DeepSeek V4.1 Flash at 552B params, foldable iPhone at $1,999, ChatGPT voice limits raised — testingcatalog · 2026-09-10
- CoT similarity test suggests Qwen3.8 may have been trained on GPT 5.5 reasoning traces — Chromix_ · 2026-09-10
- DeepSeek accused of 'pretending linear attention doesn't exist' in new architecture — teortaxesTex · 2026-09-10
- Hands-on with GLM 5.3 Flash: great at coding and research, poor at trading and ideation — ManagementNo5153 · 2026-09-10
- Tell an agent it has 1M token budget and it reasons 3-5x longer: a metacognition experiment — paraschopra · 2026-09-10
- DeepSeek's new open model beats GLM 5.3 and Kimi K3 at 4-10x lower price — deedydas · 2026-09-10