NeurIPS 2026 paper 'Follow the Winners': imitate winning trajectories instead of PPO/GRPO
lawrennd · x · 2026-10-07
'Follow the Winners' by Zhenwen Dai et al. was accepted at NeurIPS 2026. It offers a simple alternative to PPO and GRPO for RL post-training: draw trajectories from a replay buffer, keep the winners, and imitate them — no critic, no group rollouts. Co-authors include Joery de Vries and Neil Lawrence.
More from Research
- HAIPS@COLM 2026 workshop on human-centered LM privacy and security opens call for papers — tianshi_li · 2026-10-07
- Podcast: a distinctive meaning makes sentences memorable, new language memory research — GretaTuckute · 2026-10-07
- CUAWright: Terminal-Only Computer-Use Agent Beats GUI Harnesses, Cuts Cost 37.5% — ysu_nlp · 2026-10-07
- AI's Top 10 papers list: Rulin Shao lands two first-author picks — ShayneRedford · 2026-10-07
- Paradigm evals its math model across 7 hard benchmarks, releases full eval suite — tensorqt · 2026-10-07
- Paradigm: post-training gains hinge on combining procedural and LLM-based synthetic data — tensorqt · 2026-10-07