Training away its own bad reasoning lifts Qwen3-4B math scores by 7.5 points

imjustnewatai · x · 2026-09-13

A September 10 preprint takes an counterintuitive approach: instead of imitating an answer-fed teacher, it prompts a Qwen copy to reason carelessly, then trains the model away from those amplified patterns — with correct answers removed from training data. On Qwen3-4B, average accuracy across 7 math benchmarks rose 7.5 points vs just 1.0 from the teacher-based method; with thinking enabled the new method still gained 3 points. The author highlights the self-feedback loop: a model can generate useful improvement signals without knowing the correct solution.

Original post →

More from Research

Research channel →