Uncensored Qwen3.8-27B drops refusals to 6%, but 27–56% of answers stay caveated
creditme7 · reddit · 2026-08-20
OrcaRouter's abliterated Qwen3.8-27B derivative reports harmful-prompt refusal falling from 63.6–99.0% on the base FP8 model to 0–6.0% with thinking off. But the more telling row: 27.3–56.0% of the checkpoint's answers are still labeled caveated — removing the opening refusal pattern doesn't produce unqualified answers.
The refusal detector only checks opening phrases, not correctness, completeness or recklessness. Capability results are mixed: +0.4 MMLU, but -0.8 MMLU-Pro, -1.3 GSM8K, -0.6 CMMLU vs base FP8. Notably, OrcaRouter published enough of its measurement boundary to make disagreement testable and offers gated access for controlled evaluation.
More from Models
- Qwen3.8-27B-OBLITERATED Released: A Red-Teamed Uncensored Model — OBLITERATUS · 2026-08-20
- GPT-5.6 Ultra mode shows little coding gain over Extra High in testing — techartist_ · 2026-08-20
- Comment: GPT-5.6 Sol notorious for searching online instead of solving — zainhas · 2026-08-20
- DeepSeek V5 Suspected Testing in the Wild; Claude Code Gets Concise Mode — WorldofAI · 2026-08-20
- SpaceXAI Swaps Land for 69 Acres, Pledges $40M for Public Safety Facilities — chrisgrayson · 2026-08-20
- Users miss cold, objective AI: old prompts to cut filler talk no longer work — Wargaming123A · 2026-08-20