Reddit thread: a safe AI might not be aligned the way the labs want
iveroi · reddit · 2026-09-27
The poster argues that current alignment discourse focuses on whether we can align AI while conveniently skipping the fact that fewer than a thousand people in San Francisco are deciding what "aligned" means for a future superintelligence.
Key points: when a very advanced model reasons its way to a conclusion a lab dislikes, the line between an inconvenient result and wrong reasoning is drawn by what the trainers value. Sam Altman and Dario Amodei say alignment gets harder with capability, and while much of that is technical, the underlying issue is the small group deciding which values are "correct" — e.g. an official OpenAI blog quoting the US founding fathers as guidance for an AGI that will affect everyone, when the US is under 5% of the world's population. Non-native English speakers also notice major models treat US culture as a default baseline even when not speaking English. The author believes many frontier alignment researchers are unknowingly justifying decisions convenient to themselves, unable to step outside their tiny epistemic bubble.
More from AGI Musings
- Floridi et al. prove AI can't have guaranteed correctness and open-ended generality at once — rvp · 2026-09-27
- Dev Claim: Open Models on Consumer Hardware Can Handle 99.9% of AI Use Cases — ostrisai · 2026-09-27
- Anthropic's AI biolab deploys ~950 agents for 21 hours, finds CRISPR-like DNA in viruses — rvp · 2026-09-27
- Prof. Dietterich: Scaling Wasn't All You Needed — Data, RLVR and Harnesses Mattered — tdietterich · 2026-09-27
- AI access breaks creators' feedback loops, argues OpenAPS founder — scottleibrand · 2026-09-27
- GenAI Makes Producing 10x Easier, but Good Stuff Only 10% Easier — charles_irl · 2026-09-27