AI safety efforts backfired: from Anthropic to RLHF, well-meaning guardrails sped the race up

erikphoel · x · 2026-09-28

A thought-provoking thread argues that the AI world's most prominent safety efforts may have each accelerated the very race they aimed to slow down:

The author believes these individuals and organizations were sincere, but their efforts spectacularly backfired and made the world less safe — and fears much of society's ongoing work to control AI will repeat the pattern.

Original post →

More from AGI Musings

AGI Musings channel →