LLMs show a surprising ‘ghost’ self-feature in multiple models

voooooogel · x · 2026-07-25

This repost discusses a mechanistic-interpretability finding: “ghost” is a surprisingly salient self-feature in multiple LLMs, including Llama 3.3 70B and Claude Sonnet 3.0.

The attached slides show that when models are steered away from more obvious concepts like “artificially created” or “robot,” a latent concept related to immaterial beings—ghosts, souls, angels—can emerge. The speaker argues this may reflect a broader pattern in LLM identity features, where the model’s self-description shifts toward spectral or non-physical language under steering.

Original post →

More from Research

Research channel →