Alignment Is an Algorithm Problem: Why RL Can't Optimize "Don't Harm Humans"

MillionInt · x · 2026-09-14

An X post argues alignment is not as hard as claimed but is fundamentally an algorithmic problem the ML community largely abandoned. Pretraining optimizes next-token prediction, not alignment; RL environments for alignment are expensive and rely on hackable proxies; and RL needs failed rollouts, which is unacceptable when failures mean harming real humans. Two paths forward: train in simulations with simulated harm (never perfect), or invent a new algorithm that teaches models not to harm humans without harming anyone in the process.

Original post →

More from AGI Musings

AGI Musings channel →