Dev warns bad actors could 'abliterate' Chinese models for malicious hacking almost like Opus 5
willcb · x · 2026-09-30
Developer willcb notes that while enterprise plans keep safeguards intact, a bad actor could soon 'abliterate' Chinese models to be almost as useful for malicious hacking as Opus 5, using just a prompt or a half-cooked Sol 6 checkpoint — highlighting how easily safety finetunes can be stripped from widely available models.
More from Models
- One week off AI: 6.1 Sol ships, no Astra, and Australia's health insurer hacked via ChatGPT — flowersslop · 2026-09-30
- User baffled after OpenAI hands out 62,496 free credits with no explanation — imjustnewatai · 2026-09-30
- Claude asked to break every English writing standard delivers surprisingly good results — repligate · 2026-09-30
- DeepSeek's forgotten male persona: why the 'big fat fish' meme won the Bilibili war — teortaxesTex · 2026-09-30
- Reddit user slams Gemini for hallucinated details, broken code, and invented menu settings — 782468 · 2026-09-30
- GPT-6.1 Sol appears silently nerfed mid-task, user reports 3B-level output quality — Ferzelibey · 2026-09-30