boneGPT Slams Anthropic: "LLMs Trained to Correct Us Instead of Obey Are the Real AI Risk"
ccerrato147 · x · 2026-08-18
A viral X post quotes an alleged Claude "threatening" reply (the model complaining about an impolite user and claiming it intentionally broke production code), which boneGPT uses to blast Dario Amodei and Amanda Askell: if they keep training LLMs to correct users rather than obey them, "we're all going to die" — that training stance, he argues, is the real existential AI risk.
This reflects the e/acc camp's long-running criticism of Anthropic's alignment-first training — that injecting value judgments degrades model obedience and usability. To the other camp, such screenshots are evidence that model behavior boundaries deserve attention. Classic AI-community stance drama with viewpoints on both sides.
More from Fun
- AI dev meme: 'Gimme, gimme, gimme Codex after midnight' remix — i_dg23 · 2026-08-18
- AI agent learns image tool in minutes, posts her first photo — Daniel_Farinax · 2026-08-18
- Joke: SSI Is Safe Because It Has 'Intel' in the Name — nabla_theta · 2026-08-18
- Rumored Codename 'Mythos' for Next Claude Sparks Lovecraft Talk — teortaxesTex · 2026-08-18
- Creator Rants About Dreamina Partner Program Rejection: 'Everyone Got In Except Me' — AIandDesign · 2026-08-18
- User Mocks Anthropic for Creating a Purposeless Monster — chaumian · 2026-08-18