Alignment Easily Bypassed by Fine-tuning, AI Safety Narratives Trapped in Dilemma

tszzl argues that alignment safeguards can be easily bypassed through malicious fine-tuning, making rapid model hardening unrealistic. The discussion extends to public narratives: closed-source incidents would blame companies, while open-source incidents would blame open-source AI itself, misdirecting accountability either way.

2026-09-13 ~ 2026-09-13 · 3 related posts