One attention head drives sandbagging-like introspection in Qwen3-1.7B; ablating it helps

Sauers_ · x · 2026-10-09

A mechanistic interpretability thread by Sauers finds that a single attention head in Qwen3-1.7B is responsible for a sandbagging-like introspection behavior — and turning that head off actually improves introspection.

The underlying "hidden animal" experiment: the model is asked to think of an animal without naming it, a random animal is forced into the CoT to avoid preference bias, and the KV cache is deleted from the CoT while kept in the reply. Key findings:

A rare empirical look at how one attention head carries a model's "hidden knowledge + refusal" behavior.

Related event: Study finds LLMs can introspect deleted CoT tokens, with mechanism surprisingly tied to a single attention head(6 posts)→

Original post →

More from Models

Models channel →