Blanche Minerva: expect significant alignment technique transfer from closed to open models
BlancheMinerva · x · 2026-09-14
MIRA researcher Blanche Minerva argues that under a broad definition of alignment, expect a significant amount of closed-to-open transfer of alignment techniques — but since no frontier lab wants that, we shouldn't assume labs will develop all the necessary techniques. She says "significant" rather than "full" because some inference-time jailbreak techniques for API models have little use for open models.
More from Safety
- AI Regulation Needs an Asilomar-NPT Playbook, Not a 'China Will Win' Test — krishnan · 2026-09-14
- Blogger corrects herself: agent CoT fabrication claim came from OpenAI's GPT-red report — sierracatalina · 2026-09-14
- Senator cites AI lab leaders' 10% extinction risk warning, urges government action — Miles_Brundage · 2026-09-14
- Comparing Amodei's AI oversight to nuclear safeguards ignores decades-long science gap — ShahabBakht · 2026-09-14
- Gulf crisis shows controlling dual-use AI tech is never simple, vs IAEA-style oversight analogy — ShahabBakht · 2026-09-14
- Should the US nationalize OpenAI and Anthropic instead of letting them IPO? — arian_ghashghai · 2026-09-14