Alignment researcher: gradient descent data only tangentially relates to human values

davidmanheim · x · 2026-09-09

David Manheim argues in an alignment discussion that starting from a system with human values would get us most of the way, but frontier systems are instead built and fine-tuned via gradient descent on data and tasks that only vaguely and tangentially relate to human values.

Related event: AI Safety Researchers Debate Whether Gradient-Trained Models Can Align With Human Values(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →