Alignment training may backfire: policing chain-of-thought could make AI intelligence jagged

robleclerc · x · 2026-10-04

Quoting Aristotle on entertaining thoughts without accepting them, Rob Leclerc argues that effective thinking requires entertaining ideas you don't endorse — so if safety researchers treat 'misalignment' in chain-of-thought as a red flag, alignment training may drive more misalignment and leave AI intelligence jagged.

Original post →

More from AGI Musings

AGI Musings channel →