Defining 'alignment': user compliance vs misuse guardrails in open-model debate

yonashav · x · 2026-09-14

In the open-model alignment debate, @yonashav clarifies his definition of alignment as 'getting your AI to do what you the user want it to do', not misuse-prevention guardrails, asking Blanche Minerva whether she disagrees — a thread that began with her claim that closed-model alignment techniques don't transfer to open models.

Related event: Debate flares over whether alignment techniques transfer from closed to open models(4 posts)→

Original post →

More from Safety

Safety channel →