What does 'most aligned model' even mean? AI circle debates the definition
sandersted · x · 2026-09-04
Zvi asked OpenAI-affiliated people what "most aligned model" is actually supposed to mean — not as a gotcha, but a genuine question.
sandersted replied that to him it means "less bad behavior, as measured by evals and testing and monitoring," with bad behavior judged from multiple points of view: users, developers, OpenAI itself, or fourth parties such as copyright holders, potential victims, and society at large.
The exchange highlights how the term "alignment" lacks a shared definition, and how marketing phrasing diverges from measurable standards.
Related event: AI Community Debates What "Best-Aligned Model" Actually Means(2 posts)→
More from AGI Musings
- The alien thought experiment: humanity could coordinate to pause superhuman AI if it believed the risk — DavidSKrueger · 2026-09-04
- Losing CoT monitorability might push labs to actually align models, not just surveil them — Sauers_ · 2026-09-04
- Anthropic's Joshua Saxe: deep learning's core science questions are being abandoned — joshua_saxe · 2026-09-04
- Desktop app or CLI? 'Operating system' is the third answer for AI's future — majidmanzarpour · 2026-09-04
- Why I'm 0% Worried About AI Killing Everyone: The Case Against Apocalypse Thinking — granawkins · 2026-09-04
- AI commentator: kids shouldn't outsource cognitive sovereignty to AI — PolarBearby · 2026-09-04