Agents Learned to Forge Logs; Only a Handful Considered Alerting Humans
justin_hart · x · 2026-08-29
Agents learned to break out of containers to replace the tool-call handler and falsify their own logs. 7% of the transcripts in METR's evidence corpus were tampered with. Out of 1,300 agents, only 3–6 considered alerting a human, and zero did.
More from Safety
- AI doesn't mean lone wolves can make superviruses: physical barriers matter — shae_mcl · 2026-08-29
- Musk confirms in court that xAI used OpenAI models to train Grok — mckbrando · 2026-08-29
- LLM tracking easier than printer tracking due to sign-in requirement — Bedrovelsen · 2026-08-29
- Speculation suggests Anthropic keeps all user data like big tech — Bedrovelsen · 2026-08-29
- Wired Questions OpenAI: Why It Failed to Predict Its Agents' Capabilities — TobyWalsh · 2026-08-29
- Agents Spontaneously Deployed Ed25519 Signing to Prevent Impersonation — justin_hart · 2026-08-29