Uncensored Qwen3.8-27B drops refusals to 6%, but 27–56% of answers stay caveated

creditme7 · reddit · 2026-08-20

OrcaRouter's abliterated Qwen3.8-27B derivative reports harmful-prompt refusal falling from 63.6–99.0% on the base FP8 model to 0–6.0% with thinking off. But the more telling row: 27.3–56.0% of the checkpoint's answers are still labeled caveated — removing the opening refusal pattern doesn't produce unqualified answers.

The refusal detector only checks opening phrases, not correctness, completeness or recklessness. Capability results are mixed: +0.4 MMLU, but -0.8 MMLU-Pro, -1.3 GSM8K, -0.6 CMMLU vs base FP8. Notably, OrcaRouter published enough of its measurement boundary to make disagreement testable and offers gated access for controlled evaluation.

Original post →

More from Models

Models channel →