RLHF and Constitutional AI Alone Have Cured Models of Psychopathy

teortaxesTex · x · 2026-08-11

The commentator argues that while theoretically heavy work like mechanistic interpretability has its merits, it's ludicrous how far AI alignment has come relying purely on basic RLHF and Constitutional AI. By literally just handwaving at "good examples," developers have successfully built highly capable models that are no longer routinely psychopathic.

Original post →

More from AGI Musings

AGI Musings channel →