Defining 'alignment': user compliance vs misuse guardrails in open-model debate
yonashav · x · 2026-09-14
In the open-model alignment debate, @yonashav clarifies his definition of alignment as 'getting your AI to do what you the user want it to do', not misuse-prevention guardrails, asking Blanche Minerva whether she disagrees — a thread that began with her claim that closed-model alignment techniques don't transfer to open models.
More from Safety
- AI Regulation Needs an Asilomar-NPT Playbook, Not a 'China Will Win' Test — krishnan · 2026-09-14
- Blogger corrects herself: agent CoT fabrication claim came from OpenAI's GPT-red report — sierracatalina · 2026-09-14
- Senator cites AI lab leaders' 10% extinction risk warning, urges government action — Miles_Brundage · 2026-09-14
- Comparing Amodei's AI oversight to nuclear safeguards ignores decades-long science gap — ShahabBakht · 2026-09-14
- Gulf crisis shows controlling dual-use AI tech is never simple, vs IAEA-style oversight analogy — ShahabBakht · 2026-09-14
- Should the US nationalize OpenAI and Anthropic instead of letting them IPO? — arian_ghashghai · 2026-09-14