Correction: the SmolVLM data-filtering trick is from DeepSeek's earlier tech report
eliebakouch · x · 2026-09-10
Elie Bakouch clarifies his earlier post: the meticulous pretraining data handling — using SmolVLM to score image-text quality and extract high-quality interleaved data — comes from DeepSeek's earlier tech report.
Related event: DeepSeek Found Using SmolVLM to Curate Pretraining Image-Text Data(3 posts)→
More from Research
- Meta's Boxer at ECCV: closing the 3D ground-truth gap with 2D scaling — ducha_aiki · 2026-09-10
- TRL ships 1M-token long-context training guide, trains Qwen3-8B on one 8-GPU node — QGallouedec · 2026-09-10
- Hugging Face shows how to train on 1M-token sequences on a single 8-GPU node — QGallouedec · 2026-09-10
- SyncWorld turns world models into zero-shot robot simulators via visual calibration — Yuncong Yang · 2026-09-10
- Feng Yao wins ECVA PhD Award at ECCV 2026 for 3D humans + language thesis — Michael_J_Black · 2026-09-10
- Devs call for standardized "model performance across harnesses" evals — zainhas · 2026-09-10