jd_pressman on alignment: a seed of caring in models is worth tracking

jd_pressman · x · 2026-09-12

jdpressman argues against a mental shortcut in alignment debates: dismissing a model's observed kind behavior because it "might not care under further optimization." If there's a seed of caring in the model now, he argues, you should pay attention to it and figure out how to amplify it rather than predict it away.

Related event: Researcher Critiques Reductively Treating AI Goodwill as Hidden Malice(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →