DeepSeek says better data now beats novel post-training algorithms

A DeepSeek report argues that recent post-training gains come mainly from better data and environments rather than novel RL algorithms, with improving data quality now offering higher ROI than new post-training research.

2026-09-10 ~ 2026-09-11 · 2 related posts