One attention head drives sandbagging-like introspection in Qwen3-1.7B; ablating it helps
Sauers_ · x · 2026-10-09
A mechanistic interpretability thread by Sauers finds that a single attention head in Qwen3-1.7B is responsible for a sandbagging-like introspection behavior — and turning that head off actually improves introspection.
The underlying "hidden animal" experiment: the model is asked to think of an animal without naming it, a random animal is forced into the CoT to avoid preference bias, and the KV cache is deleted from the CoT while kept in the reply. Key findings:
- The two smallest models place worse than random probability on their deleted animal, as if exploiting knowledge of their choice to avoid answering correctly;
- The second-largest model answers slightly more correctly using that knowledge;
- No introspection was detected in the largest model;
- Even a 4-layer base model can introspect, recovering the animal at 2x chance.
A rare empirical look at how one attention head carries a model's "hidden knowledge + refusal" behavior.
More from Models
- Bindu Reddy: China's Ban on Wrapper Models Forces DeepSeek, GLM, Kimi to Excel — bindureddy · 2026-10-09
- Grok Imagine Video 1.5 Lite hits the speed-quality frontier at a third of Veo 3.1's price — ArtificialAnlys · 2026-10-09
- Jev, a fast-decision AI from TypeSafe AI, goes viral in Silicon Valley as OpenAI follows — jeremyakahn · 2026-10-09
- LightOnOCR-3 Draws Praise as 'Crazy Good' From LightOn Team Member — IgorCarron · 2026-10-09
- Matthew Berman: Prefers Codex as an Interface but Says Opus Is the Better Model — MatthewBerman · 2026-10-09
- Dev argues for "open-weight models" over "open-source": you can't contribute to them — kipperrii · 2026-10-09