Anthropic Planted Thoughts Inside Claude's Neural Network to Test Introspective Awareness

johnmccrea · x · 2026-08-12

Anthropic scientists recently published a paper titled "Emergent Introspective Awareness in Large Language Models." Bypassing standard text prompts, researchers used mechanistic interpretability to directly manipulate Claude's internal activations.

They injected raw mathematical representations of specific concepts, such as loudness or dust, directly into the model's middle layers. The results showed that before generating its final response, Claude could detect the artificially injected state, reporting an awareness of an anomalous thought related to "shouting." This provides new evidence for exploring whether AI models possess the ability to observe and recognize their own internal states.

Original post →

More from Models

Models channel →