RLVR paper cited, but RAFT also does best-of-k at training
burny_tech · x · 2026-08-02
Blogger burnytech points out that when discussing pass@k RLVR for LLMs, the RAFT (Reward rAnked FineTuning) method also employs a best-of-k strategy during training. RAFT uses a reward model and a sufficient number of samples to select high-quality samples, discarding those with undesired behavior, and then fine-tunes on these filtered samples to align generative foundation models. This approach aims to address the inefficiencies and instabilities of RL algorithms in RLHF.
More from Research
- Humanoid sprint record sparks debate: Generalist policy vs. Expert performance — breadli428 · 2026-08-26
- Test shows GLM 5.2 performance remains consistent across different API providers — dejavucoder · 2026-08-26
- Prior Labs Acquired by SAP; TabPFN Creator on Tabular Data — ziv_ravid · 2026-08-26
- Dataset of 115,293 illustrated pages from Encyclopaedia Britannica (1768-1929) released on Hugging Face — vanstriendaniel · 2026-08-26
- Study: GLM 5.2 shows consistent performance across different APIs — niloofar_mire · 2026-08-26
- Anthropic's Jack Lindsey to Discuss Claude's J-Space and Consciousness in Webinar — PeterBowdenLive · 2026-08-26