DeepSeek says data quality ROI now beats novel post-training algorithms like GRPO follow-ups
yacineMTB · x · 2026-09-12
- According to a cited note, DeepSeek now believes improving data quality delivers far higher ROI than designing new post-training algorithms, despite having invented GRPO, iterating it after R1, and adding MOPD.
- Contrast: ByteDance, Alibaba, Google, and Meta publish heavily on post-training algorithms yet underperform their GPU endowments.
- @lujasper argues this has long been true outside labs: 80% of post-training effort should go into data — hiring experts to audit RL tasks, manually sifting rollouts and SFT data for suspicious samples, ensuring tasks are passable, and keeping data diverse in difficulty and category.
Related event: DeepSeek Says Data Quality Now Beats New Post-Training Algorithms(3 posts)→
More from Infra
- Intel engineers show wafer-level chiplet testing for co-packaged optics paper — jwt0625 · 2026-09-12
- Intel engineers spotted testing chiplets and packages on wafer-level tester — jwt0625 · 2026-09-12
- Ex-Googler speculates Apple could build a next-token predictor native to Apple Silicon — yaroslavvb · 2026-09-12
- Yaroslav Bulatov launches A100-targeted architecture experiments attacking backprop's memory wall — yaroslavvb · 2026-09-12
- ADSP Ep. 303: Mark Saroufim on Open vs Closed Models, AI GPU Kernels and Autoresearch — blelbach · 2026-09-12
- Repurpose an old low-VRAM GPU just for mmproj in llama.cpp — an order of magnitude faster — inthesearchof · 2026-09-12