Qwen Uncensored Model Sparks Concern: Risks of Removing Safety Rails
gregpr07 · x · 2026-08-19
User tests Qwen 3.8 Uncensored, noting it performs any web request without safety gates. This raises concerns about 'abliterated' models, highlighting that despite heavy investment in safety by Anthropic and OpenAI, safety rails can be easily stripped via RLHF, emphasizing the fragility and necessity of safety research.
More from Safety
- Debate erupts over lethal military robots vs. failing civilian units — teortaxesTex · 2026-08-24
- Only 1 of 20 Potential Presidential Candidates Answered AI Pause Query — DavidSKrueger · 2026-08-24
- Chinese Transforming Robot Dog Sparks US Trade Policy Criticism — TinfoilTricorn · 2026-08-24
- Turkey blocks at least 12 Grok posts on national security grounds — Unusual_Variation293 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24
- RSI concepts: automated development vs. capability acceleration — fleetingbits · 2026-08-24