Don't Ask Models to 'Solve Alignment' Whole — Decompose It Into Goodhart and Value Loading, Researcher Argues

jd_pressman · x · 2026-09-09

jdpressman offered a more nuanced stance in the alignment debate: while vaguely asking frontier models to "solve alignment" is risky because it points them at the "dumb Yuddite crap" region of latent space, breaking the problem into parts like "value loading", "causal and extremal Goodhart", and "value extrapolation" makes it no more dangerous than any other math problem.

The reply counters quetzalrainbow's claim that asking an agent swarm to literally solve alignment is the worst idea ever, highlighting a genuine disagreement in the alignment community over whether frontier models should work on alignment research itself.

Related event: AI researchers debate whether agent swarms should 'solve alignment'(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →