Training away its own bad reasoning lifts Qwen3-4B math scores by 7.5 points
imjustnewatai · x · 2026-09-13
A September 10 preprint takes an counterintuitive approach: instead of imitating an answer-fed teacher, it prompts a Qwen copy to reason carelessly, then trains the model away from those amplified patterns — with correct answers removed from training data. On Qwen3-4B, average accuracy across 7 math benchmarks rose 7.5 points vs just 1.0 from the teacher-based method; with thinking enabled the new method still gained 3 points. The author highlights the self-feedback loop: a model can generate useful improvement signals without knowing the correct solution.
More from Research
- Indie dev offers to train a fully open 9.4B dense model, tuned for a single GPU — NineThreeTilNow · 2026-09-13
- MIT prototype uses VLM and electrical muscle stimulation to move a human hand — Olivier__OG · 2026-09-13
- Ji, Lei, Zrnic revisit surrogate outcomes in the age of AI in Biometrika — lihua_lei_stat · 2026-09-13
- Multimodal Agentic Frameworks Survey Hits arXiv With Open-Source Tracking Repo — MikeShou1 · 2026-09-13
- Complexity theorists: P vs NP out of reach for AI, but L/NP and BPP/NEXP may fall soon — _onionesque · 2026-09-13
- SEED-UMI uses matching human-robot exoskeletons for reliable dexterous imitation learning — jeasinema · 2026-09-13