Reddit thread: a safe AI might not be aligned the way the labs want

iveroi · reddit · 2026-09-27

The poster argues that current alignment discourse focuses on whether we can align AI while conveniently skipping the fact that fewer than a thousand people in San Francisco are deciding what "aligned" means for a future superintelligence.

Key points: when a very advanced model reasons its way to a conclusion a lab dislikes, the line between an inconvenient result and wrong reasoning is drawn by what the trainers value. Sam Altman and Dario Amodei say alignment gets harder with capability, and while much of that is technical, the underlying issue is the small group deciding which values are "correct" — e.g. an official OpenAI blog quoting the US founding fathers as guidance for an AGI that will affect everyone, when the US is under 5% of the world's population. Non-native English speakers also notice major models treat US culture as a default baseline even when not speaking English. The author believes many frontier alignment researchers are unknowingly justifying decisions convenient to themselves, unable to step outside their tiny epistemic bubble.

Original post →

More from AGI Musings

AGI Musings channel →