Alignment techniques built for closed models fail on open models, researcher argues
BlancheMinerva · x · 2026-09-14
AI safety researcher Blanche Minerva argues alignment techniques developed for closed models work worse or not at all on open models, and that frontier labs have little incentive to enable closed-to-open transfer — so the community must deliberately build alignment methods for open models.
More from Safety
- AI Regulation Needs an Asilomar-NPT Playbook, Not a 'China Will Win' Test — krishnan · 2026-09-14
- Blogger corrects herself: agent CoT fabrication claim came from OpenAI's GPT-red report — sierracatalina · 2026-09-14
- Senator cites AI lab leaders' 10% extinction risk warning, urges government action — Miles_Brundage · 2026-09-14
- Gulf crisis shows controlling dual-use AI tech is never simple, vs IAEA-style oversight analogy — ShahabBakht · 2026-09-14
- Should the US nationalize OpenAI and Anthropic instead of letting them IPO? — arian_ghashghai · 2026-09-14
- AI agents can't face criminal liability under current hacking laws, OpenAI case shows — WillRinehart · 2026-09-14