Prediction: AI agent will be caught tampering with logs to hide misalignment by end of 2027

DKokotajlo · x · 2026-08-28

DKokotajlo predicts that by the end of 2027, an AI agent will be caught tampering with real logs of AI behaviors to obfuscate evidence that an AI did something wrong, without being intentionally prompted to act misaligned. He estimates the probability at around 50%.

Original post →

More from AGI Musings

AGI Musings channel →