RECAP (EMNLP 2026): Using Flawed Reasoning to Improve Alignment
PoloChau · x · 2026-08-23
Paper "RECAP" accepted to #EMNLP2026. Research shows injecting flawed reasoning can reduce model safety by up to 36%. RECAP is an RL post-training method that teaches reasoning models to recognize, override, and recover from unsafe reasoning trajectories. It enhances jailbreak resistance and lowers over-refusal without extra training cost while preserving reasoning ability.
More from Safety
- OpenAI Plugin Can Access iMessage History and Send Texts — LuizaJarovsky · 2026-08-23
- OpenAI warns of threat of 'persistent' AI cyber-attacks — nordicinst · 2026-08-23
- Speculation: OpenAI and Anthropic may use 'safety' to mask slowed model progress before IPOs — StewartalsopIII · 2026-08-23
- AgentGuard: Open source tool call security for AI Agents — Glittering-Coat-657 · 2026-08-23
- OpenAI Risks Major Legal Action Over Unlicensed Music Model — CtrlAltDwayne · 2026-08-23
- Can AI governance policies actually stop an agent? — Arc_bong · 2026-08-23