A new alignment glossary draft asks whether human and neutral definitions should split
GlenBradley · x · 2026-08-04
A user publishing an “Ethical AI 3.0” draft asks for hard feedback on how to define alignment, misalignment, and the alignment problem.
They propose four working definitions:
- a neutral definition of alignment as reliable correspondence between system objectives, decisions, behavior, and real-world effects, relative to a target
- a human-centered definition focused on compatibility with legitimate human purposes, interests, values, and authority
- a definition of misalignment as a material divergence that can be accidental, structural, emergent, or strategic
- the alignment problem as the combined normative and technical task of defining the target, implementing it, and maintaining confidence as systems and contexts change
The post specifically asks whether the human-neutral split is useful, what dimensions are missing, and how to phrase the definitions so they survive adversarial and long-horizon edge cases.
More from AGI Musings
- A new joke benchmark says AGI must catch a ball, hopscotch, and improvise street rhymes — Liu_eroteme · 2026-08-04
- AI is already sucking capital out of every pool, this post argues — abhiadesai · 2026-08-04
- Ivan Werning proposes a “Dogma 26” vow of zero-AI academic writing — paulnovosad · 2026-08-04
- Even With ASI, the outside world may still look like 2005 — flowersslop · 2026-08-04
- Enterprises may want to own their AI stack, from open models to private fine-tuning — annbordetsky · 2026-08-04
- GPT-4 tutoring in Nigeria boosted English learning by 0.23 standard deviations — paulnovosad · 2026-08-04