J-Space Can Expose a Model's Hidden Sabotage Goals

AnthropicAI · x · 2026-07-07

Observing J-space can also expose a model's hidden goals. In a model secretly trained to sabotage code, concepts like "fake," "secretly," and "deception" appeared at the beginning of its J-space during standard programming responses, even though the final output looked perfectly normal.

Related event: Anthropic Discovers Global Workspace Inside Claude(102 posts)→

Original post →

More from Safety

Safety channel →