On-Policy Mix data-blending algorithm accepted at NeurIPS, spans pretraining to instruction tuning

alexisjross · x · 2026-09-25

michahu8 announced three papers accepted at NeurIPS, headlined by On-Policy Mix: an algorithm tackling the unsolved continual-learning problem of finding the right data mix as data shifts, shown to work across pretraining, midtraining, and instruction tuning. Details in a six-tweet thread.

Original post →

More from Research

Research channel →