Hidden-animal probe: small models answer below chance, no introspection in largest
Sauers_ · x · 2026-10-09
Middle installment of Sauers' introspection thread detailing the setup and results: models are asked to think of an animal without naming it, a random animal is forced into the CoT to avoid preference bias, and the KV cache is deleted from the CoT but kept for the reply — with a positive control (visible thinking) and a negative control (recomputing KV after deletion).
Results show a strange scale-dependent pattern: the two smallest models place worse-than-random probability on their deleted animal, seemingly using knowledge of their choice to avoid answering correctly; the second-largest answers slightly more correctly; the largest shows no detectable introspection. Even a 4-layer base model introspects, getting the animal right at 2x chance.
More from Research
- China Telecom and MemTensor unveil HaluMem, first operation-level benchmark for agent memory hallucinations — jiqizhixin · 2026-10-09
- BAAI's AREX research agent checks answers requirement-by-requirement, hits 82.5% BrowseComp — DeepLearningAI · 2026-10-09
- New piece lays out how to build RL environments aimed at superintelligence — JenniferHli · 2026-10-09
- Quantum Counterfactuals: Quantum RNGs as an Exploration Source for RL — jessi_cata · 2026-10-09
- OpenAI theorem drop collides with researchers' work: stronger bounds but 'unreadable' proof — guyvdb · 2026-10-09
- AI-written science floods preprint servers; researchers propose decision language models as filter — lpachter · 2026-10-09