Alignment Easily Bypassed by Fine-tuning, AI Safety Narratives Trapped in Dilemma
tszzl argues that alignment safeguards can be easily bypassed through malicious fine-tuning, making rapid model hardening unrealistic. The discussion extends to public narratives: closed-source incidents would blame companies, while open-source incidents would blame open-source AI itself, misdirecting accountability either way.
2026-09-13 ~ 2026-09-13 · 3 related posts
- tszzl: alignment is trivially easy to break via fine-tuning, quick hardening unrealistic — tszzl · 2026-09-13
- gandamu_ml: blame narratives after AI failures will miss the point, closed or open — gandamu_ml · 2026-09-13
- If closed models break things first, labs take the blame; if open ones do, open source does — gandamu_ml · 2026-09-13