tszzl: alignment is trivially easy to break via fine-tuning, quick hardening unrealistic
tszzl · x · 2026-09-13
Reacting to @yonashav's optimism, tszzl argues models can't be hardened everywhere quickly, and alignment is trivially easy to break if someone fine-tunes a model for malicious purposes. He stresses he is making a prediction, not advocacy, and hopes he is wrong.
More from AGI Musings
- 'AI bubble bursting': CEO slowdown messaging may trigger Monday crash, user warns — SumitGup · 2026-09-13
- Scobleizer dissects AI engagement bait: a thread claiming Altman, Amodei and Musk agreed to slow frontier models — Scobleizer · 2026-09-13
- AI companies born from safety fears keep 'succumbing to capitalism', thread argues — birchlse · 2026-09-13
- What domain expertise still can't be trusted to Claude or ChatGPT? — sartomiki · 2026-09-13
- Circuit complexity is stuck — and AI could be the perfect adversarial partner to crack it — _onionesque · 2026-09-13
- OpenAI Researcher Warns AI-Driven Military Power Will Concentrate in Frontier Labs — jachiam0 · 2026-09-13