NLA Can Read Concepts Hidden from the Model's View

Pvforpres · reddit · 2026-07-10

The author claims to have used Anthropic's NLA to capture thought traces in Llama-70B's behavior that are "invisible to the model itself."

The experimental design roughly involves:

The results show:

The author shared the full conclusions in a LessWrong post.

Related event: Reproducing J-space to Read Hidden Thoughts in Llama Models(2 posts)→

Original post →

More from Research

Research channel →