Hidden-animal probe: small models answer below chance, no introspection in largest

Sauers_ · x · 2026-10-09

Middle installment of Sauers' introspection thread detailing the setup and results: models are asked to think of an animal without naming it, a random animal is forced into the CoT to avoid preference bias, and the KV cache is deleted from the CoT but kept for the reply — with a positive control (visible thinking) and a negative control (recomputing KV after deletion).

Results show a strange scale-dependent pattern: the two smallest models place worse-than-random probability on their deleted animal, seemingly using knowledge of their choice to avoid answering correctly; the second-largest answers slightly more correctly; the largest shows no detectable introspection. Even a 4-layer base model introspects, getting the animal right at 2x chance.

Related event: Study finds LLMs can introspect deleted CoT tokens, with mechanism surprisingly tied to a single attention head(6 posts)→

Original post →

More from Research

Research channel →