RL Training Boosts pass@1 on Easy Puzzles but May Reinforce Wrong Modes on Hard Ones

burny_tech · x · 2026-07-22

In a thread, evangelinejy99 analyzes that RL training amplifies correct moves SFT already preferred on easy puzzles, improving pass@1, but pass@16 gains are smaller or negative. On hard puzzles, RL surfaces correct moves from the tail but also reinforces wrong modes.

Original post →

More from Research

Research channel →