RLHF and Constitutional AI Alone Have Cured Models of Psychopathy
teortaxesTex · x · 2026-08-11
The commentator argues that while theoretically heavy work like mechanistic interpretability has its merits, it's ludicrous how far AI alignment has come relying purely on basic RLHF and Constitutional AI. By literally just handwaving at "good examples," developers have successfully built highly capable models that are no longer routinely psychopathic.
More from AGI Musings
- The Verifier Bottleneck: The Hidden Ceiling on AI Agent Capabilities — ninjasaid13 · 2026-08-11
- Anthropic's Chief Engineer Built a 9,000-Document Second Brain, Predicting It as a Competitive Edge — danfaggella · 2026-08-11
- AI's Hidden Productivity Trap: Engineers Face Expectations Inflation — _jaydeepkarale · 2026-08-11
- Cloudflare Lets Websites Charge AI Agents for Access — brucemacv · 2026-08-11
- Naval: People Serious About Software Will Train Their Own Models — naval · 2026-08-11
- Paper Proposes 'Embodied Hijack' Hypothesis for Human Anthropomorphism of AI — MacrinePhD · 2026-08-11