Analysis of V4-0731 and V4-0813 Training Methods and Performance

teortaxesTex · x · 2026-08-16

The tweet analyzes the differences between model versions V4-0731 and V4-0813. V4-0731 was baked with a mix of popular agents like Pi/Codex and serves as the main RL teacher. V4-0813 was distilled from Flash, focusing on the minimal harness setting to finalize the DSH project, explaining the variance in their reasoning overfitting.

Related event: DeepSeek Models 0731 vs 0813 Spark Generalization Debate(2 posts)→

Original post →

More from Models

Models channel →