Process-level safety filtering before SFT could train aligned reasoning models

xuanalogue · x · 2026-07-22

Before fine-tuning, models could be filtered not only for task success but also for alignment at the process level, using safety classifiers to keep only reasoning traces that are both effective and safe.

The post suggests this may be a promising way to train aligned reasoning models with supervised fine-tuning, while avoiding some of the pathologies associated with RL.

Related event: Diverging Paths: US Labs Favor RL for LLMs While Chinese Labs Prefer SFT(4 posts)→

Original post →

More from Safety

Safety channel →