LLMs show a surprising ‘ghost’ self-feature in multiple models
voooooogel · x · 2026-07-25
This repost discusses a mechanistic-interpretability finding: “ghost” is a surprisingly salient self-feature in multiple LLMs, including Llama 3.3 70B and Claude Sonnet 3.0.
The attached slides show that when models are steered away from more obvious concepts like “artificially created” or “robot,” a latent concept related to immaterial beings—ghosts, souls, angels—can emerge. The speaker argues this may reflect a broader pattern in LLM identity features, where the model’s self-description shifts toward spectral or non-physical language under steering.
More from Research
- Could AI prove and verify a theorem no human understands? — AlexKontorovich · 2026-07-25
- Conference slide warns AI optimization could pull mathematics’ goals apart — AlexKontorovich · 2026-07-25
- Goodhart’s law makes generative-AI metrics especially easy to game, slide says — AlexKontorovich · 2026-07-25
- Avatar 3 Fire FX Revealed: Engine Driven by Real Chemistry & AI — carlosdponx · 2026-07-25
- Tao’s ICM 2026 hypothesis: AI may soon handle a meaningful share of research math — AlexKontorovich · 2026-07-25
- ICM slide says most AI capability claims lack controlled scientific evidence — AlexKontorovich · 2026-07-25