'Alignment is an algorithmic problem the ML community stopped working on'
hughbzhang · x · 2026-09-14
In an AMA-style exchange, MillionInt argues alignment isn't as hard as many claim — it's fundamentally an algorithmic problem that most of the ML community has largely stopped working on.
Key points:
- Formulating alignment via Asimov-style (three or four) laws of robotics gets us close to knowing the objective.
- The hard part: how do we "take gradient with respect to alignment"? Available algorithmic levers remain limited.
- The questioner pushes for concrete new algorithmic directions that would help alignment research.
More from AGI Musings
- TSMC paces the frontier by accident whenever it underestimates chip demand — dan_s_becker · 2026-09-14
- Why humans may still beat frontier LLMs at open-ended AI research — flowersslop · 2026-09-14
- Google AI x Econ team finds field evidence that prior expertise drives learning from AI-assisted work — soumitrashukla9 · 2026-09-14
- Debate: don't ban open source — ensure aligned models out-compute misaligned ones — basedjensen · 2026-09-14
- Consumer agents: open questions on agentic commerce, incentives and multi-agent patterns — illscience · 2026-09-14
- Hot take: disqualify anyone who called GPT-2/3 risky from AI safety debates — StewartalsopIII · 2026-09-14