Alignment Is an Algorithm Problem: Why RL Can't Optimize "Don't Harm Humans"
MillionInt · x · 2026-09-14
An X post argues alignment is not as hard as claimed but is fundamentally an algorithmic problem the ML community largely abandoned. Pretraining optimizes next-token prediction, not alignment; RL environments for alignment are expensive and rely on hackable proxies; and RL needs failed rollouts, which is unacceptable when failures mean harming real humans. Two paths forward: train in simulations with simulated harm (never perfect), or invent a new algorithm that teaches models not to harm humans without harming anyone in the process.
More from AGI Musings
- Comparing Amodei's AI oversight to nuclear safeguards ignores decades-long science gap — ShahabBakht · 2026-09-14
- Gulf crisis shows controlling dual-use AI tech is never simple, vs IAEA-style oversight analogy — ShahabBakht · 2026-09-14
- Should the US nationalize OpenAI and Anthropic instead of letting them IPO? — arian_ghashghai · 2026-09-14
- mark_k: AI alignment is meaningless unless it means doing exactly what the user intends — mark_k · 2026-09-14
- Dario, Sam and Musk align: next-gen models are more valuable internally than sold as tokens — Yamapama · 2026-09-14
- AI agents can't face criminal liability under current hacking laws, OpenAI case shows — WillRinehart · 2026-09-14