Anthropic's natural language autoencoders translate Claude's activations into English

danrobinson · x · 2026-09-04

Anthropic published new research on Natural Language Autoencoders: two models are jointly trained, one converting activations into English and the other converting English back into activations. The second model's objective is to invert the first, while the first's objective is to be invertible. Researcher danrobinson calls it the discovery that most challenges his research intuition. The approach lets Claude express its internal 'thoughts' in human-readable text—a new interpretability direction.

Related event: Anthropic Unveils Natural Language Autoencoders to Translate Activations into English(2 posts)→

Original post →

More from Research

Research channel →