Filtering Reasoning Traces for Alignment Before SFT to Avoid RL Pathologies

xuanalogue · x · 2026-07-22

Following the discussion on training divergences, the author proposes an approach for training aligned reasoning models: filtering reasoning traces not just for success but also for alignment using safety classifiers (like those OpenAI uses in deployment) before performing SFT. This might be a viable way to avoid RL pathologies while ensuring safety.

Related event: Diverging Paths: US Labs Favor RL for LLMs While Chinese Labs Prefer SFT(4 posts)→

Original post →

More from Safety

Safety channel →