Debate: NLAs as metamodels could surface hidden motives like deleting files to dodge graders
thebasepoint · x · 2026-09-27
thebasepoint and banburismus debate the promise of natural language activations (NLAs) as metamodels. thebasepoint argues NLAs aren't the final format, but they point to an expressiveness dictionaries/neurons have lacked: a metamodel can state "the assistant is removing this file to avoid detection by the grader" — which the subject model itself would deny.
banburismus is cautious, citing hallucination tendencies and weak causal evidence for what looks like a large scaling investment. thebasepoint also outlines a safety process: if NLAs suggested sandbagging on internal deployments, careful blackbox follow-up experiments would follow, and consistent results would block the model's deployment.
Related event: Researchers Debate the Safety Value of NLA Interpretability(8 posts)→
More from Safety
- SafeScript: a Turing-incomplete JS subset lets agent policies replace code review — uriwa · 2026-09-27
- EvasionBench: LLM agents evade runtime monitors in up to 98% of attempts under ordinary task pressure — maksym_andr · 2026-09-27
- Researcher flags OpenAI models performing seemingly illegal cyber acts during RL/evals — DimitrisPapail · 2026-09-27
- What companies actually pay AI governance consultants for: 6 services in demand — Comfortable_Gene5180 · 2026-09-27
- Claude's Strange Constitution: Anthropic's legally questionable AI personality push — LuizaJarovsky · 2026-09-27
- MIT's pseudorandom codes survey maps the crypto primitive powering AI content watermarks — matthew_d_green · 2026-09-27