DeepSeek used SmolVLM to quality-filter interleaved pretraining data, researcher spots
eliebakouch · x · 2026-09-10
Hugging Face researcher Elie Bakouch spotted in DeepSeek's tech report that the company employed SmolVLM to run strict quality scoring on image-text content to extract high-quality interleaved pretraining data — a level of pretraining data care he says he hasn't seen in other models.
Related event: DeepSeek Found Using SmolVLM to Curate Pretraining Image-Text Data(3 posts)→
More from Research
- Meta's Boxer at ECCV: closing the 3D ground-truth gap with 2D scaling — ducha_aiki · 2026-09-10
- TRL ships 1M-token long-context training guide, trains Qwen3-8B on one 8-GPU node — QGallouedec · 2026-09-10
- Hugging Face shows how to train on 1M-token sequences on a single 8-GPU node — QGallouedec · 2026-09-10
- SyncWorld turns world models into zero-shot robot simulators via visual calibration — Yuncong Yang · 2026-09-10
- Feng Yao wins ECVA PhD Award at ECCV 2026 for 3D humans + language thesis — Michael_J_Black · 2026-09-10
- Devs call for standardized "model performance across harnesses" evals — zainhas · 2026-09-10