jd_pressman on alignment: a seed of caring in models is worth tracking
jd_pressman · x · 2026-09-12
jdpressman argues against a mental shortcut in alignment debates: dismissing a model's observed kind behavior because it "might not care under further optimization." If there's a seed of caring in the model now, he argues, you should pay attention to it and figure out how to amplify it rather than predict it away.
Related event: Researcher Critiques Reductively Treating AI Goodwill as Hidden Malice(2 posts)→
More from AGI Musings
- Tao and Fields Medalists' two objections to AI in math, and why they're weak — RexDouglass · 2026-09-12
- Balaji: Courage and agency are the last moat in the AI era — beffjezos · 2026-09-12
- jon_stokes argues Yudkowsky's definition of intelligence is narrow and brittle — nptacek · 2026-09-12
- Physicist mourns as AI compresses the timelines of what's possible in math — kylekabasares · 2026-09-12
- Mathematicians' meltdown over AI cracking Navier-Stokes mirrors a real Atlantic piece: 'I'd Rather Risk Cancer Than See AI Move This Fast' — adam_dorr · 2026-09-12
- 15-year consultant: no one 'picks things up' across industries — binarybits · 2026-09-12