Scholars Criticize OpenAI's Narrow Definition of Alignment: Instruction-Following Isn't Enough

davidmanheim · x · 2026-08-09

A researcher highlighted that OpenAI's definition of "alignment" appears overly narrow, focusing primarily on models "following their spec or obeying instructions."

The critique argues that an AI blindly obeying instructions written by a small group is, by default, misaligned with humanity at large. Furthermore, no single spec can cover all out-of-distribution cases, which contradicts how applied ethics actually works. The safety community is urged to address this mismatch in problem definition.

Original post →

More from AGI Musings

AGI Musings channel →