tszzl: alignment is trivially easy to break via fine-tuning, quick hardening unrealistic

tszzl · x · 2026-09-13

Reacting to @yonashav's optimism, tszzl argues models can't be hardened everywhere quickly, and alignment is trivially easy to break if someone fine-tunes a model for malicious purposes. He stresses he is making a prediction, not advocacy, and hopes he is wrong.

Related event: Alignment Easily Bypassed by Fine-tuning, AI Safety Narratives Trapped in Dilemma(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →