RLVR paper cited, but RAFT also does best-of-k at training

burny_tech · x · 2026-08-02

Blogger burnytech points out that when discussing pass@k RLVR for LLMs, the RAFT (Reward rAnked FineTuning) method also employs a best-of-k strategy during training. RAFT uses a reward model and a sufficient number of samples to select high-quality samples, discarding those with undesired behavior, and then fine-tunes on these filtered samples to align generative foundation models. This approach aims to address the inefficiencies and instabilities of RL algorithms in RLHF.

Original post →

More from Research

Research channel →