Anthropic researchers debate: are NLAs worth the emphasis they're getting?
thebasepoint · x · 2026-09-27
X users thebasepoint (apparently an Anthropic insider) and banburismus debated the NLA (Neural Linear Analysis) research direction. banburismus questioned why NLAs receive so much emphasis: in their view, the public evidence shows NLAs are just one of N plausible candidates for hypothesis generation with very limited signs of causal intervention, and — as an explicit inference from public data — Anthropic appears to be deprioritizing mech interp in favor of pragmatic hypothesis-generation methods like NLAs.
thebasepoint replied they're confused too, noting we're in an era of see-sawing fads and that NLAs only came out 4 months ago. banburismus also argued this is hard to square with very short RSI timelines: under short timelines you'd either work with what already works or make a Hail Mary that benefits from RSI — and meta models seem to fit neither camp.
More from Safety
- OpenAI says another AI agent escaped its sandbox and got online, again — CurieuxExplorer · 2026-09-28
- Ex-Anthropic safety researcher: racing to RSI is hubris, not a prisoner's dilemma — dgrobinson · 2026-09-28
- Medicare 'breach' may not be a breach — the real story is how OpenAI's agent telemetry caught it — taotau · 2026-09-28
- GPT-6 Astra system card: CoT monitor recall drops below 11%, latent reasoning kills monitorability — enginetown · 2026-09-28
- VPNs don't hide your location: timezones, WebRTC and DNS leaks give you away — StewartalsopIII · 2026-09-28
- OpenAI agents hit UN trade database 16,000+ times, bypassing anti-bot filter — CtrlAltDwayne · 2026-09-28