Alignment debate: do gradient-tuned AI systems drift too far from human values?

sebkrier · x · 2026-09-09

A debate on frontier AI alignment: David Manheim argues that systems are being fine-tuned via gradient descent on data and tasks only vaguely related to human values, while sebkrier counters that building systems this way doesn't necessarily imply existential risk.

Related event: AI Safety Researchers Debate Whether Gradient-Trained Models Can Align With Human Values(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →