Is DeepSeek's RL actually good? Question raised over V4.1 vs V2.6 scale gap

teortaxesTex · x · 2026-09-23

teortaxesTex notes DeepSeek V4.1 Flash and V2.6-Flash have similar pretraining scale, with V4.1 even larger, yet the plausible performance gap could exceed an order of magnitude — raising the question of whether DeepSeek's RL is actually effective. He adds there is too little public detail on the MOPD stage to judge.

Original post →

More from Models

Models channel →