If NLAs show sandbagging on internal deployments, blackbox experiments would block deployment

thebasepoint · x · 2026-09-27

Continuing an NLA discussion thread, thebasepoint outlines a concrete safety process: if natural language activation (NLA) metamodels indicated a model was sandbagging on internal deployments, the team would run careful blackbox follow-up experiments — and if results were consistent with that explanation, deployment of the model would be blocked.

Related event: AI safety researchers debate interpretability paths and the value of NLA(8 posts)→

Original post →

More from Safety

Safety channel →