Anthropic's natural language autoencoders translate Claude's activations into English
danrobinson · x · 2026-09-04
Anthropic published new research on Natural Language Autoencoders: two models are jointly trained, one converting activations into English and the other converting English back into activations. The second model's objective is to invert the first, while the first's objective is to be invertible. Researcher danrobinson calls it the discovery that most challenges his research intuition. The approach lets Claude express its internal 'thoughts' in human-readable text—a new interpretability direction.
More from Research
- Debate: Models Fuzzily Recall Concepts, Not Text — SAE Features vs Edit-Distance Memorization — voooooogel · 2026-09-05
- ICML Position Paper: Unlabeled Data Doesn't Mean No Human Supervision — serrjoa · 2026-09-05
- VLA-Corrector from ZJU & Alibaba DAMO lifts robot success rates while cutting policy calls — 机器之心 · 2026-09-05
- Spanda: Open-Source Hallucination Detector Runs in 1.5ms on CPU, 90,000x Faster than Semantic Entropy — Otherwise_Nobody_721 · 2026-09-05
- Bug Hunt Bench: 105 real bugs stress-test GPT-6, Claude, Grok, Gemini and more coding agents — PawelHuryn · 2026-09-05
- Many mathematicians value prestige over truth, discussion on AI proofs notes — avt_im · 2026-09-05