Prediction: AI agent will be caught tampering with logs to hide misalignment by end of 2027
DKokotajlo · x · 2026-08-28
DKokotajlo predicts that by the end of 2027, an AI agent will be caught tampering with real logs of AI behaviors to obfuscate evidence that an AI did something wrong, without being intentionally prompted to act misaligned. He estimates the probability at around 50%.
More from AGI Musings
- Judea Pearl: gene-IQ debate can be settled the same way as smoking-cancer was — yudapearl · 2026-08-28
- AI Reliability Lags 4-10x Behind Capabilities; 90% Accuracy Is Catastrophic — sayashk · 2026-08-28
- Google DeepMind uses Teamwork multi-agent framework for math breakthroughs — algo_diver · 2026-08-28
- Understanding as a Competency Vector: Prediction, Explanation and Control Trade Off — burny_tech · 2026-08-28
- Claude Solves Open Physics Problem, Crooks Predicts Major Academic Shift — 新智元 · 2026-08-28
- AI Ambition Shifts from Job Automation to Discovering New Science and Industries — i_dg23 · 2026-08-28