RECAP (EMNLP 2026): Using Flawed Reasoning to Improve Alignment

PoloChau · x · 2026-08-23

Paper "RECAP" accepted to #EMNLP2026. Research shows injecting flawed reasoning can reduce model safety by up to 36%. RECAP is an RL post-training method that teaches reasoning models to recognize, override, and recover from unsafe reasoning trajectories. It enhances jailbreak resistance and lowers over-refusal without extra training cost while preserving reasoning ability.

Original post →

More from Safety

Safety channel →