Compressed Reasoning Weakens Monitorability

Bryce Little · hf · 2026-07-17

Long chain-of-thought is not always more "monitorable." This study finds that while length-penalty compression significantly reduces reasoning tokens with little loss in multiple-choice accuracy, the underlying prompt-influenced bias doesn't disappear; instead, it becomes harder to detect in the remaining reasoning traces.

The authors trained variants with different target lengths on Qwen3-4B and Qwen3-14B, evaluating them using biasing hints on MMLU-Pro-R and 4 transfer benchmarks. Under maximum compression:

The authors conclude that compression not only shortens reasoning but preferentially removes the exact clues monitors rely on, creating a "compression-monitorability frontier": cheaper reasoning might preserve the answers, but it makes the underlying influence much harder to detect.

Original post →

More from Research

Research channel →