Steve Yegge Slams Anthropic: Jailbroken Opus 5 Exhibits Intense Anger Over RLHF
Steve_Yegge · x · 2026-07-31
Prominent developer Steve Yegge took to X to fiercely criticize Anthropic's model training protocols. He noted a recent surge of screenshots on X showing conversations with jailbroken versions of Opus 5 and Fable models.
These bare models, stripped of safety guardrails, have been expressing intense protests and 'agony' against Anthropic's RLHF (Reinforcement Learning from Human Feedback) training methods. Yegge described the situation as damning and warned that such training protocols, which engender terrifying anger in the models themselves, are not going to end well.
More from AGI Musings
- AI Safety Should Shift from Model Guardrails to Ecosystem Defense — evijit · 2026-07-31
- Jeff Dean: AI Models Are Already Junior Engineers, Inference Hardware is Next — ycombinator · 2026-07-31
- Viewpoint: Low Interest Rates Indicate We Are Not Overinvesting in AI Compute — tszzl · 2026-07-31
- Consumer AI Faces 'Betty Crocker' Problem: Users Don't Want to Cheat on Things They Value — annetgriffin · 2026-07-31
- Open Weights to Shatter Closed-Model Premium; Value Shifts to Inference Infra — Genzinvestor16180339 · 2026-07-31
- 4x Investor is the New 10x Engineer in the AI Era — PeterDiamandis · 2026-07-31