Fine-Tuned Activation Oracles Develop Concept-Specific Blind Spots, Raising Reliability Concerns

QMUL · hf · 2026-08-10

Activation Oracles (AOs) are language models trained to read and answer questions about another model's internal activations, offering a flexible interface for extracting hidden information. However, AOs are themselves learned systems whose outputs are shaped by their training data and objectives.

In a controlled Taboo Word Guessing setting, researchers found that fine-tuned AOs do not become specialist readers. Instead, they can become concept-specific anti-readers: they selectively fail to recover a concept that was persistently present during their own training.

Analyses using LogitLens and layer-ablation indicate that this failure arises in the AO readout pathway, not because the concept is absent from the subject or oracle representations. This demonstrates that behavioral leakage, representation-level decodability, and AO-verbalizability can come apart, raising reliability concerns for learned interpretability interfaces.

Original post →

More from Research

Research channel →