rgblong clarifies his AI safety critique isn't Roko's basilisk reasoning
rgblong · x · 2026-09-18
Responding to Nina Panickssery's jab that his take resembles Roko's basilisk reasoning, rgblong clarifies it doesn't: he does want alternatives to Anthropic's alignment approach and shares some safety worries about it. He adds that his concern centers on the safety risks of instilling contradictory or incoherent views about consciousness, goals, and self into models.
More from AGI Musings
- tszzl: the sci-fi taboo against synthetic life is emerging as a real-world force — tszzl · 2026-09-18
- EA Movement's Arc: "Give Everything to Charity" Then, Rule the AI World Now — wordgrammer · 2026-09-18
- Agent swarms are the third scaling axis: 10,000+ agents behind recent model breakthroughs — paraschopra · 2026-09-18
- Why Anthropic's Repligate held a vigil, not a funeral, for retiring Claude models — repligate · 2026-09-18
- 1988 sci-fi short story imagined an emergent Chinese Room that answered back — toptickcrypto · 2026-09-18
- OpenAI's Noam Brown: Air-gapping may not stop misaligned AI, safety bar must rise — basedjensen · 2026-09-18