Why Capabilities May Be Easier Than Alignment: Agents Correct Beliefs but Protect Goals

gleech · x · 2026-09-07

gleech articulates the live hypothesis that "capabilities are easier than alignment": a coherent agent corrects its own mistaken beliefs because they hurt its goals, and protects its goals because revising them usually harms them. Combined with the argument that more diverse data eventually fixes capabilities but never fixes what a model wants — goals can stay behaviorally indistinguishable even in the limit — this sketches why value understanding is capability-like but not value loading.

Related event: gleech argues alignment lags capabilities: capability errors surface, goal errors only show after harm(9 posts)→

Original post →

More from Safety

Safety channel →