Training on 69,000 DeepSeek-R1 Failed Attempts Boosts 7B Math Scores

imjustnewatai · x · 2026-08-06

A recent paper suggests that a strong model's failed trajectories can be more valuable for training than its successes. Researchers collected 69,000 failed math reasoning attempts from DeepSeek-R1 and trained smaller models to locate the mistakes, repair the reasoning, and solve the problems.

Experiments showed that a 7B model trained this way achieved significant math improvements:

The core insight is that a strong model's wrong answer typically contains long stretches of correct reasoning followed by a single critical error. Teaching a model to locate and recover from that mistake provides more learning signal than copying a polished answer. This implies that the massive amounts of failed coding runs and agent trajectories stored in AI labs hold immense untapped training value.

Original post →

More from Research

Research channel →