Finding an Inner Monologue Representation Space in LLMs
mlpowered · x · 2026-07-07
The mlpowered team identified a small subset of representations inside language models, dubbed J-space. Functioning like the model's "inner monologue," it presents intermediate concepts, makes judgments about the current text, and can even detect prompt injection. This offers new clues for model interpretability and safety.
Related event: Anthropic Discovers Global Workspace Inside Claude(102 posts)→
More from Research
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11