The Self-Distill-Zero Method

prfsanjeevarora · x · 2026-07-10

Self-Distill-Zero is a simple self-improvement training pipeline for scenarios involving reward or correctness discriminators. The model first generates answers, then uses discriminator scores for online distillation, providing denser token-level supervision from the model to itself.

The post states that this method significantly outperforms SFT and RL baselines on math and coding tasks. It also mentions that the paper won Best Paper at the ICML RLxF workshop and received an Honorable Mention at the AI4Math workshop.

Original post →

More from Research

Research channel →