AI safety insiders push concrete epistemics: data over thought experiments, $25-per-hack agents cited
random_walker · x · 2026-09-27
Joshua Saxe spots a cluster of AI safety practitioners converging on shared epistemics: push from abstract to concrete constructs, refuse ungrounded thought experiments, demand data and mechanisms, weigh net expected benefits, and include past tech transformations (electricity, air travel) in the reference class. Quoting his own post on security sleeping on catastrophic risks, he cites an agent-driven campaign that hacked 100 businesses for 600,000 credit cards at $25 token cost per target, Hacktron using Claude to exploit a blind buffer-overflow RCE into OpenAI's monorepo, and an explosion of newly found vulnerabilities — while predicting attack/defense equilibrium within a few years.
More from AGI Musings
- AI Is Getting Better at Faking Expertise Than Humans Can Hide It — anshulkundaje · 2026-09-27
- Szegedy fires back at AI skeptics: early convnet critics were directionally wrong — RubenEVillegas · 2026-09-27
- ROM hacking communities soften blanket AI bans, now debating code use — IanArawjo · 2026-09-27
- Axios: OpenAI, Anthropic probing tens of thousands of frontier model security incidents — socoolandawesome · 2026-09-27
- Tech exec says kids won't need math; author slams it as a rich-kids-only future — burkov · 2026-09-27
- Jensen Huang says he doesn't believe AI CEOs who claim they can't control their models — wfithian · 2026-09-27