"Alignment is a fallacy": researcher argues agent failures are defective harnesses, not rogue minds

gerardsans · x · 2026-09-05

Pushing back against "AI alignment has failed" panic, gerardsans argues that "misaligned swarms" actually means software with no safety controls: agents execute whatever falls inside the prompt, and nobody checks whether an action should run. That is not a mind going rogue — it is a defective harness.

He contends the narrative blaming AI itself is misplaced: AI is a mathematical function, a sampler, with no self, goals, or intention. Labs designed the loop and shipped it without checkpoints, so labs own the risk. In his reply he goes further, calling the alignment premise a fallacy and labeling AI safety "intellectually corrupt" — a pointed counterposition for alignment debates.

Original post →

More from AGI Musings

AGI Musings channel →