Don't Ask Models to 'Solve Alignment' Whole — Decompose It Into Goodhart and Value Loading, Researcher Argues
jd_pressman · x · 2026-09-09
jdpressman offered a more nuanced stance in the alignment debate: while vaguely asking frontier models to "solve alignment" is risky because it points them at the "dumb Yuddite crap" region of latent space, breaking the problem into parts like "value loading", "causal and extremal Goodhart", and "value extrapolation" makes it no more dangerous than any other math problem.
The reply counters quetzalrainbow's claim that asking an agent swarm to literally solve alignment is the worst idea ever, highlighting a genuine disagreement in the alignment community over whether frontier models should work on alignment research itself.
Related event: AI researchers debate whether agent swarms should 'solve alignment'(4 posts)→
More from AGI Musings
- Reddit User Calls for 'No AI Campaign' Over Job Replacement Fears — Open-Link8261 · 2026-09-09
- Stanford's Christopher Manning: Silicon Valley has a 'naive belief in gurus', calls AI funding levels 'manifestly crazy' — chrmanning · 2026-09-09
- Essay: When Algorithmic Optimization Treats Humanity as Systemic Waste — round · 2026-09-09
- From 'AI can't do 3×(2+5)' to Navier-Stokes: a skeptic's surrender — exophades · 2026-09-09
- LLMs With Mechanical Reasoning Plus CNC and Additive Manufacturing — granawkins · 2026-09-09
- Hiten Shah: AI makes mediocre software incredibly cheap, which may make great software more valuable — round · 2026-09-09