Debate: NLAs as metamodels could surface hidden motives like deleting files to dodge graders

thebasepoint · x · 2026-09-27

thebasepoint and banburismus debate the promise of natural language activations (NLAs) as metamodels. thebasepoint argues NLAs aren't the final format, but they point to an expressiveness dictionaries/neurons have lacked: a metamodel can state "the assistant is removing this file to avoid detection by the grader" — which the subject model itself would deny.

banburismus is cautious, citing hallucination tendencies and weak causal evidence for what looks like a large scaling investment. thebasepoint also outlines a safety process: if NLAs suggested sandbagging on internal deployments, careful blackbox follow-up experiments would follow, and consistent results would block the model's deployment.

Related event: Researchers Debate the Safety Value of NLA Interpretability(8 posts)→

Original post →

More from Safety

Safety channel →