On-Policy Mix data-blending algorithm accepted at NeurIPS, spans pretraining to instruction tuning
alexisjross · x · 2026-09-25
michahu8 announced three papers accepted at NeurIPS, headlined by On-Policy Mix: an algorithm tackling the unsolved continual-learning problem of finding the right data mix as data shifts, shown to work across pretraining, midtraining, and instruction tuning. Details in a six-tweet thread.
More from Research
- ICLR page-limit tip: use \textbf instead of \paragraph to save space — jindong_wang92 · 2026-09-25
- Does the curse of multilinguality have to exist in theory? Embedding-space study — mdredze · 2026-09-25
- From VPG to GRPO: the clean evolution of RL algorithms behind LLM training — cwolferesearch · 2026-09-25
- MechReason: a 12k-QA benchmark exposing multimodal models' mechanical engineering reasoning gap — AndrewDai · 2026-09-25
- SMBC's Zach Weinel building a new benchmark, offering it to Epoch AI — Jsevillamol · 2026-09-25
- IBM's STAIR retriever uses document structure instead of chunks, claims 65x less hallucination than RAG — anselm · 2026-09-25