Anthropic Unveils Natural Language Autoencoders to Translate Activations into English

Anthropic has introduced Natural Language Autoencoders, jointly training two models—one that translates internal activations into English and another that maps English back to activations. Paradigm researcher Dan Robinson called the surprising results enough to upend his research intuitions.

2026-09-04 ~ 2026-09-04 · 2 related posts