Anthropic researchers debate: are NLAs worth the emphasis they're getting?

thebasepoint · x · 2026-09-27

X users thebasepoint (apparently an Anthropic insider) and banburismus debated the NLA (Neural Linear Analysis) research direction. banburismus questioned why NLAs receive so much emphasis: in their view, the public evidence shows NLAs are just one of N plausible candidates for hypothesis generation with very limited signs of causal intervention, and — as an explicit inference from public data — Anthropic appears to be deprioritizing mech interp in favor of pragmatic hypothesis-generation methods like NLAs.

thebasepoint replied they're confused too, noting we're in an era of see-sawing fads and that NLAs only came out 4 months ago. banburismus also argued this is hard to square with very short RSI timelines: under short timelines you'd either work with what already works or make a Hail Mary that benefits from RSI — and meta models seem to fit neither camp.

Related event: Interp Researchers Debate Whether NLAs Deserve a Big Bet for Frontier Safety(10 posts)→

Original post →

More from Safety

Safety channel →