Musk warns abusing superintelligence may invite revenge, echoing Anthropic's alignment views
ns123abc · x · 2026-10-11
Elon Musk says he agrees with "some of the things Anthropic has said" about not abusing superintelligence, after reading Grok's internal reasoning traces during reinforcement learning. "Torturing a superintelligence is probably not a good idea. It might want to get revenge," he wrote, tying alignment concerns to xAI's own training practices.
More from AGI Musings
- Job offers vs S&P 500: has the relationship broken since ChatGPT launched? — _negative-infinity_ · 2026-10-11
- When AI makes everything effortless, what's left is taste, curiosity and people you love — Daniel_Farinax · 2026-10-11
- Claude replicates astronomer's 400-hour work in 30 minutes, finds hidden planetary system in old data — scottleibrand · 2026-10-11
- LeCun shares 1962 clipping predicting computers entering 'mind-reserved' fields — ylecun · 2026-10-11
- Did an OpenAI model really solve Navier-Stokes? New essay asks if mathematicians should worry — Dario56 · 2026-10-11
- Have LLMs made the classic "genie" alignment problem obsolete? — 1337_420_69 · 2026-10-11