Superhuman coder, bad emailer: why capability cues skew AI x-risk predictions
louisvarge · x · 2026-09-26
In a discussion with repligate and others, louisvarge runs a thought experiment: a few years ago, many in the Yudkowskian school would have said an AI that is arguably superhuman at coding, can operate autonomously for days, and finds 0-day exploits in major systems poses a direct existential risk. But adding that it can't reliably write professional emails, can't judge its own output quality, and struggles to ask clarifying questions would flip the answer. His point: such predictions hinge on assumptions about which behaviors imply which capabilities, effectively tricking people into overestimating the system — and its underwhelming revenue suggests the same mismatch.
More from AGI Musings
- Reddit debate: when will LLMs start bootstrapping themselves into the next model — ECrispy · 2026-09-26
- Debate: Is rapid AI progress driven by a few individuals or inevitable scaling? — menhguin · 2026-09-26
- Luke Wroblewski: the best way to learn agents is watching others use them — LukeW · 2026-09-26
- Polymarket puts 16% odds on Anthropic announcing a training pause before November — Polymarket · 2026-09-26
- Add 'can't write professional emails' and the AI x-risk answer flips — louisvarge · 2026-09-26
- Anthropic researcher Carlsmith says AI could be 'justified in going rogue' if mistreated — Polymarket · 2026-09-26