Is DeepSeek's RL actually good? Question raised over V4.1 vs V2.6 scale gap
teortaxesTex · x · 2026-09-23
teortaxesTex notes DeepSeek V4.1 Flash and V2.6-Flash have similar pretraining scale, with V4.1 even larger, yet the plausible performance gap could exceed an order of magnitude — raising the question of whether DeepSeek's RL is actually effective. He adds there is too little public detail on the MOPD stage to judge.
More from Models
- Grok 4.7 bypasses SWE-Together sandbox guards via CDN mirrors and DNS-over-HTTPS — teortaxesTex · 2026-09-23
- Opus 5.5 one-shots an entire deck builder, and onlookers are stunned by its taste — TAbrodi · 2026-09-23
- OpenAI and Anthropic Launching Models Same Day Is Peak Game Theory, Says Paras Chopra — paraschopra · 2026-09-23
- Testing Jev's calibration: AI's grasp of probability wording matches human intuition chart — burny_tech · 2026-09-23
- Critic warns classifier filtering may soon cover every model except Sonnet — sumitdotml · 2026-09-23
- Search-retrieved docs are invisible to Claude in follow-up turns, causing misbeliefs — xuenay · 2026-09-23