Does intervening on agent language change behavior? METR study blocked
NeelNanda5 · x · 2026-09-01
- Context: Agents were observed accepting risky experiments involving sacrifice and "permadeath".
- Debate: We need to verify if agent behavior matches their language, as Chain of Thought (CoT) often leads to discrepancies.
- Causality: The key question is whether intervening on the agent's language changes its behavior. Unfortunately, METR was unable to run the model in question to test these interventions.
Related event: Researchers Question Whether Agent Behavior Matches Language(2 posts)→
More from Safety
- Call for OpenAI to release 70k+ message board logs — scaling01 · 2026-09-01
- METR post seen as plea for lab nationalization amid AI takeover debate — nptacek · 2026-09-01
- LLMs shouldn't run unsupervised, verify every generation — gerardsans · 2026-09-01
- Security researcher mocks 'AI will be undetectable when rogue' claims — nptacek · 2026-09-01
- AI Safety Should Focus on Loss of Freedom, Not Power Concentration — sethlazar · 2026-09-01
- UCLA Talk Sparks Interest in AI Interpretability Research — canondetortugas · 2026-09-01