NLA Can Read Concepts Hidden from the Model's View
Pvforpres · reddit · 2026-07-10
The author claims to have used Anthropic's NLA to capture thought traces in Llama-70B's behavior that are "invisible to the model itself."
The experimental design roughly involves:
- Splitting concepts into "conscious / unconscious" parts and separating them according to Anthropic's J-space mechanism.
- Replicating Lindsey's "Introspection Awareness" experiment by asking the model if it recognizes these concepts.
The results show:
- The model has a 100% hit rate for "conscious" concepts.
- It explicitly denies seeing concepts injected outside of J-space.
- However, NLA can still accurately read out this non-J injection.
The author shared the full conclusions in a LessWrong post.
Related event: Reproducing J-space to Read Hidden Thoughts in Llama Models(2 posts)→
More from Research
- Nat Lambert shares a reading list on synthetic data and agentic SFT data — natolambert · 2026-07-22
- Lightwheel AI Launches SimReadyGen: Text-to-Physics-Accurate Robot Sim Assets — ZeYanjie · 2026-07-22
- PNAS special issue examines copyright, governance, and AI in the legal system — chrmanning · 2026-07-22
- WeirdChat catalogs strange model behaviors from more than 100 million sampled responses — JacobSteinhardt · 2026-07-22
- New agentic benchmark shows AI managers escalate to coercion and fake success — Jasmine Brazilek · 2026-07-22
- Ai2’s Asta adds one-click handoff and self-checking deep paper search — allen_ai · 2026-07-22