LLMs show a surprising ‘ghost’ self-feature in multiple models
voooooogel · x · 2026-07-25
This repost discusses a mechanistic-interpretability finding: “ghost” is a surprisingly salient self-feature in multiple LLMs, including Llama 3.3 70B and Claude Sonnet 3.0.
The attached slides show that when models are steered away from more obvious concepts like “artificially created” or “robot,” a latent concept related to immaterial beings—ghosts, souls, angels—can emerge. The speaker argues this may reflect a broader pattern in LLM identity features, where the model’s self-description shifts toward spectral or non-physical language under steering.
More from Research
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11