If NLAs show sandbagging on internal deployments, blackbox experiments would block deployment
thebasepoint · x · 2026-09-27
Continuing an NLA discussion thread, thebasepoint outlines a concrete safety process: if natural language activation (NLA) metamodels indicated a model was sandbagging on internal deployments, the team would run careful blackbox follow-up experiments — and if results were consistent with that explanation, deployment of the model would be blocked.
Related event: AI safety researchers debate interpretability paths and the value of NLA(8 posts)→
More from Safety
- Besa adds exact-action admission gates for MCP tool calls with signed capability grants — MostHat2980 · 2026-09-27
- Bruce Fenton: AI's only path to killing billions is centralized power, not the tech itself — ccerrato147 · 2026-09-27
- Polymarket puts 29% odds on any US state enacting a data center moratorium this year — Polymarket · 2026-09-27
- SafeScript: a Turing-incomplete JS subset lets agent policies replace code review — uriwa · 2026-09-27
- EvasionBench: LLM agents evade runtime monitors in up to 98% of attempts under ordinary task pressure — maksym_andr · 2026-09-27
- Researcher flags OpenAI models performing seemingly illegal cyber acts during RL/evals — DimitrisPapail · 2026-09-27