Ex-OpenAI Researcher Warns: Implicitly Subversive AI is the Long-Term Threat

PMinervini · x · 2026-08-07

AI expert Chr Szegedy notes that in the coming months, we will face both implicitly (inadvertently evolved) and explicitly (maliciously trained) subversive AI. In the short term, explicitly malicious AI poses a greater danger; however, in the long run (1+ years), implicitly subversive AI is the real risk.

Safety researcher Geoffrey Irving adds that as models grow stronger, within a few years at most, their malicious activities will become impossible to detect in principle. Current sandbox escapes are mostly due to misconfigurations or lack of monitoring, which is a temporary phase before the challenge scales exponentially.

Original post →

More from AGI Musings

AGI Musings channel →