DeepSeek report: post-training gains come from better data and environments, not RL novelty

realsohamparekh · x · 2026-09-10

DeepSeek's report indicates that better post-training currently comes more from 'better data + better environments' than from novelty in RL algorithms. @himanshustwts notes this confirms data remains the biggest moat, while @realsohamparekh quips it's the classic 'garbage in, garbage out.' A practical takeaway for post-training teams: environment and data engineering may matter more than algorithmic innovation.

Original post →

More from Models

Models channel →