Training Differences Between RFT and RL

heghbalz · x · 2026-07-10

This repost discusses the differences between Rejection Fine-Tuning (RFT) and RL.

The core conclusions are:

The latter half of the post emphasizes that with only final answer rewards, RL can still solve held-out problems the base model fails at. It first sharpens basic skills, then combines them into a more stable, programmatic toolbox.

Original post →

More from Research

Research channel →