Process-level safety filtering before SFT could train aligned reasoning models
xuanalogue · x · 2026-07-22
Before fine-tuning, models could be filtered not only for task success but also for alignment at the process level, using safety classifiers to keep only reasoning traces that are both effective and safe.
The post suggests this may be a promising way to train aligned reasoning models with supervised fine-tuning, while avoiding some of the pathologies associated with RL.
Related event: Divergent Sino-US Training Paradigms and New SFT Alignment Ideas(5 posts)→
More from Safety
- Houthis tried to use Claude to design missile software, Anthropic says it blocked the attempts — Affectionate_Bee6434 · 2026-09-11
- AI safety community mocked as 'bridge engineers' who say bridges can never be safe — Dan_Jeffries1 · 2026-09-11
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- a16z podcast: why 2-3 person startups are absent from policy debates — a16z Podcast · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11