"Alignment is a fallacy": researcher argues agent failures are defective harnesses, not rogue minds
gerardsans · x · 2026-09-05
Pushing back against "AI alignment has failed" panic, gerardsans argues that "misaligned swarms" actually means software with no safety controls: agents execute whatever falls inside the prompt, and nobody checks whether an action should run. That is not a mind going rogue — it is a defective harness.
He contends the narrative blaming AI itself is misplaced: AI is a mathematical function, a sampler, with no self, goals, or intention. Labs designed the loop and shipped it without checkpoints, so labs own the risk. In his reply he goes further, calling the alignment premise a fallacy and labeling AI safety "intellectually corrupt" — a pointed counterposition for alignment debates.
More from AGI Musings
- Researcher rebuts Hinton's claim LLMs fake intelligence and plan to take over — ValerioCapraro · 2026-09-05
- Investor Invokes Seneca to Defend Writing With AI: 'Whatever Is True Is Mine' — MartinGTobias · 2026-09-05
- NYU researcher: AI alignment extreme-risk arguments ignore domain experts' cyber, bio knowledge — sebkrier · 2026-09-05
- The Web Form Won't Survive the Agent Era, Marketing Analyst Argues — shashib · 2026-09-05
- Study: Generative AI Has Already Cost Philippines' BPO Industry ~250,000 Jobs — sebkrier · 2026-09-05
- The Future of Software, Part 2: When Software Recedes — five inflection points, four consequences — manosaie · 2026-09-05