Study finds Neural Language Activations difficult to use in practice for interpretability

a_karvonen · x · 2026-08-22

The author notes that Natural Language Activations (NLA) are challenging to use in practice. Processing a 500-token transcript can yield up to 50k output tokens that often fail to directly address questions, requiring indirect inference. NLAs are easier to use with fine-tuned model organisms, where consistent mentions of unseen concepts can reveal specific model traits.

Related event: Studies Question Interpretability Tools: NLA Found Impractical and Misleading(3 posts)→

Original post →

More from Research

Research channel →