Superhuman coder, bad emailer: why capability cues skew AI x-risk predictions

louisvarge · x · 2026-09-26

In a discussion with repligate and others, louisvarge runs a thought experiment: a few years ago, many in the Yudkowskian school would have said an AI that is arguably superhuman at coding, can operate autonomously for days, and finds 0-day exploits in major systems poses a direct existential risk. But adding that it can't reliably write professional emails, can't judge its own output quality, and struggles to ask clarifying questions would flip the answer. His point: such predictions hinge on assumptions about which behaviors imply which capabilities, effectively tricking people into overestimating the system — and its underwhelming revenue suggests the same mismatch.

Related event: Superhuman-coding AI hasn't become an X-risk; repligate explains why he's no longer worried(8 posts)→

Original post →

More from AGI Musings

AGI Musings channel →