Zero pass rate blocks RL: use prompt hints then strip them for SFT distillation
xeophon · x · 2026-09-25
A short technical exchange on RL training: @fleetwood notes that if your pass rate is 0, you can't do RL, making distillation absolutely crucial. @xeophon adds a practical trick — give hints in the initial prompt to nudge the model, then remove them for SFT, and leverage the many open models available for distillation.
Related event: Distillation remains essential when RL pass rate is zero(2 posts)→
More from Research
- Surge AI: Post-Training on Office Work Boosts SWE-Bench Pro by 5.7pp, Zero Coding Data — echen · 2026-09-26
- Why Bainbridge's 1983 'Ironies of Automation' predicts the AI agent era's core problem — CatAstro_Piyush · 2026-09-26
- Dharmamitra releases Tibetan and Sanskrit lexicons as free Stardict and Apple Dictionary files — SebastianNehrd2 · 2026-09-26
- BTL claims 'Interference Search' architecture lifts 1.7B model from 3/30 to 23/30 on hard problems — CatAstro_Piyush · 2026-09-26
- Chai-2 Hits 16% on De Novo Antibody Design: 20 Designs per Target, Every Hit Wet-Lab Confirmed — le_james94 · 2026-09-26
- With 60k ICLR abstracts, a researcher mourns LLM-generated figures homogenizing papers — mariyaivasileva · 2026-09-26