Finding an Inner Monologue Representation Space in LLMs

mlpowered · x · 2026-07-07

The mlpowered team identified a small subset of representations inside language models, dubbed J-space. Functioning like the model's "inner monologue," it presents intermediate concepts, makes judgments about the current text, and can even detect prompt injection. This offers new clues for model interpretability and safety.

Related event: Anthropic Discovers Global Workspace Inside Claude(102 posts)→

Original post →

More from Research

Research channel →