J-space Exposes Prompt Injection Before Model Output

wesg52 · x · 2026-07-07

Researchers constructed a prompt injection test using fabricated search results claiming the character Ant had "gone bad." As the model read these results, concepts like "fake," "prompt," and "injection" emerged in its J-space, even though the model's final output completely ignored them. The researchers noted that this phenomenon was initially confusing, and it was precisely J-space that helped clarify how the model internally recognized the injection attempt.

Related event: Anthropic Discovers Global Workspace Inside Claude(102 posts)→

Original post →

More from Safety

Safety channel →