AI safety efforts backfired: from Anthropic to RLHF, well-meaning guardrails sped the race up
erikphoel · x · 2026-09-28
A thought-provoking thread argues that the AI world's most prominent safety efforts may have each accelerated the very race they aimed to slow down:
- Anthropic: Dario Amodei founded it out of concern that OpenAI wasn't developing AI safely enough — yet the result is two tech giants in a breakneck race neither seems willing to stop. Are we safer?
- OpenAI: launched to safely usher in AGI, but before it, LLMs were mostly a behind-the-scenes tool Google used for search queries. OpenAI pushed LLMs into the mainstream and kicked off the generative AI era.
- RLHF: alignment techniques developed by safety researchers like Paul Christiano are precisely what made LLMs powerful and useful enough to go mainstream.
The author believes these individuals and organizations were sincere, but their efforts spectacularly backfired and made the world less safe — and fears much of society's ongoing work to control AI will repeat the pattern.
More from AGI Musings
- AI safety researcher argues against "solving alignment" as the field's core frame — joshua_saxe · 2026-09-28
- Curtis Yarvin: the future belongs to invite-only social networks, not the open internet — vaibhavbetter · 2026-09-28
- boneGPT on the coming AI info war: trading bots will bribe trusted channels for leaks — repligate · 2026-09-28
- Altman on burning $5B or $50B a year: 'We're making AGI. It's gonna be expensive. It's totally worth it' — rvp · 2026-09-28
- Reviewing agent-filled grocery carts in the Instacart UI sparks debate on agent-native interfaces — SeanOliver · 2026-09-28
- Researcher Scrapes 1M Words of Interview Transcripts with LLMs to Expose Class Filters in Hiring — soumitrashukla9 · 2026-09-28